logoalt Hacker News

nonethewiser • yesterday at 8:30 PM • 3 replies • view on HN

How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.


Replies

fancyfredbot • today at 4:09 PM

You are Google, training a new model. You regularly benchmark the checkpoints. It's gradually improving as you train. When do you stop training and release the model?

You certainly don't do it when you are just behind the frontier. Instead you wait until you are ahead on a lot of benchmarks.

So, if Google release a new frontier model, the chances it's ahead on most benchmarks are statistically about 100% because if it isn't yet then they won't release it.

sebzim4500 • yesterday at 8:55 PM

Part cherry-picking of benchmarks, part leapfrogging

MoreThanMe • yesterday at 10:38 PM

[flagged]