logoalt Hacker News

WASDxyesterday at 4:34 PM1 replyview on HN

DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.


Replies

kimjune01yesterday at 11:45 PM

it would be nice if these benchmark reports actually specified which tasks they passed and which ones they didn't.