logoalt Hacker News

PunchyHamster • yesterday at 7:59 PM • 1 reply • view on HN

> Today's frontier models can't even handle 100% of the SWE benchmarks that have been around for longer than they've been training the models. Companies that are benchmaxxing their models against the benchmarks haven't even been able to get them to 100%.

Is average developer 100% the benchmarks ?


Replies

criemen • yesterday at 9:07 PM

Given how flawed the benchmarks are (see for example https://epoch.ai/benchmarks/deepswe/review), not even the mythical 10x developer is gonna get 100% on these benchmarks.