logoalt Hacker News

rafram • yesterday at 2:53 AM • 2 replies • view on HN

This is a pretty obscure and in-the-weeds benchmark, but to me the models’ interpretation feels quite reasonable.


Replies

deprave • yesterday at 4:53 PM

Apologies, I didn’t mean to imply it’s a benchmark, I just wanted to provide a reproducible example of where I see models make decisions that seem to be fine initially but might paint the software architecture into a challenging corner. I don’t expect models to read my mind, but I do see them produce a lot of verbose output, none of which is used to say “here’s a simple response to your ask, but have you also considered...”

shimman • yesterday at 3:59 AM

It's obscure to use common functions from the standard library?

➕ show 2 replies