logoalt Hacker News

ouz-atoday at 7:39 AM3 repliesview on HN

I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.


Replies

r_leetoday at 9:27 AM

That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?

puszczyktoday at 8:39 AM

This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing

show 1 reply
thomasnowheretoday at 7:43 AM

same here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.