I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.
This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing
same here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.
That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?