logoalt Hacker News

rjtcyesterday at 10:26 PM4 repliesview on HN

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:

https://artificialanalysis.ai/models

Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?


Replies

jrflotoday at 4:06 AM

Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.

jesse_dot_idtoday at 4:21 AM

They're gaming benchmarks HTH

aniviacatyesterday at 11:28 PM

This benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.

Bjorkbattoday at 12:19 AM

The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.

And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.

show 1 reply