logoalt Hacker News

notatoadtoday at 1:46 AM1 replyview on HN

when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.


Replies

caconym_today at 2:07 AM

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)

show 2 replies