Benchmarks are so difficult with ai because as soon as one gets popular it enters the dataset so the next iteration of the model is trained on the solution. I’m not sure if there’s any potential work around here or anyone doing interesting work but would love to hear about it if so
But if this effect were really that strong, the models should be getting nearly 100% on the common benchmarks. But for many of the benchmarks, even after being public for more than a year, the new models only get 60-80%.