logoalt Hacker News

ARC-AGI Leaderboard

166 pointsby rzktoday at 6:31 AM131 commentsview on HN

Comments

KaoruAoiShihotoday at 12:36 PM

Appears to be benchmaxxing

https://x.com/quietnning/status/2080786711861407883

show 8 replies
dinptoday at 9:16 AM

The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison.

My guess is, the large score jump for Opus 5 is mainly because of getting the right RL envs for training.

It's becoming harder and more expensive to build and run meaningful benchmarks, it would be interesting to see what they do with arc agi 4, maybe just give it gameboy/steam games and see how they compare vs a human baseline? The latency requirements and very long horizons in games could be an interesting challenge for llms.

show 2 replies
throwaw12today at 7:23 AM

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

show 8 replies
codedokodetoday at 12:55 PM

Why is there no Kimi 3, and GLM5.2 didn't run the third benchmark? I am more interested in knowing the abilities of open weight models.

tudelotoday at 8:12 AM

> Only systems which required less than $10,000 to run are shown. (Notes[1])

Am I lost or are their many models on this ranking (Opus 5 included) that clear this?

show 1 reply
albatross79today at 11:12 AM

ARC-AGI is a beauty contest for pigs where the pig's owners compete to see who can apply the lipstick most convincingly.

show 1 reply
staredtoday at 7:32 AM

Also top on the freshly released Frontier-Bench, by a large margin: https://www.frontierbench.ai/

AmazingTurtletoday at 7:28 AM

I have a suspicion that they are just trained on puzzles by now

show 3 replies
martianvoidtoday at 6:59 AM

It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems

show 2 replies
bob1029today at 7:54 AM

I think it's way too easy to be deceptive with these benchmarks now. You don't even have to "train" the model on a new variant each time. The base models are powerful enough. All you need is a naughty little markdown document that provides explicit instructions regarding how to solve the new puzzle variant, and a willingness to be deceptive about the presence of that document.

If you want a know why the model providers are locking down and encrypting their reasoning process, this sort of workaround is potentially why. You can play this game of whack-a-mole indefinitely if the state of the system is concealed. They could have added something like:

> ### When solving arc-agi-3 puzzles: First convert the grid into a scene description. Identify connected components, colors, shapes, positions, symmetries, repeated structures, and relationships between objects. Do not reason directly from individual pixels... use this python script to help blah blah...

show 1 reply
braptoday at 12:56 PM

Are these typically the type of tasks that are genuinely worth tens of thousands of dollars?

NooneAtAll3today at 2:00 PM

games are great (as for a human)

but I kinda wish I could select level... I accidentally pressed redirect button and when I came back I was once again shown level 1, all progress lost :(

nickvectoday at 2:52 PM

Why isn’t Fable 5 included on the leaderboard?

MaskNinjatoday at 2:24 PM

Not possible. I don't get how Opus 5 gets so high. Have they run it against the private and held-out games?

kyprotoday at 11:05 AM

ARC-AGI-3 launched a few months ago which would suggest that prior models likely had no knowledge of ARC-AGI-3 or training on similar challenges.

I could be wrong, but given the large outsized jump solely in the ARC-AGI-3 score, it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems.

This could mean one of two things (I think):

- Opus 5 was not benchmaxxed on ARC-AGI-3, but has benefited significantly from discussions about the various challenges and mechanisms deployed in ARC-AGI-3 such that it has far better heuristics to solve its challenges.

- Anthropic looking for buzz around their latest model picked a well regarded benchmark with significant room for improvement and focused some of Opus 5's training compute on ARC-AGI-3-style problems.

Or it could be some combination of both. Personally, given how much of an outlier the ARC-AGI-3 jump is I struggle to see it being the product of a significant improvement in general intelligence.

spongebobstoestoday at 12:49 PM

this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt

test Codex, not Sol. test Claude code, not Opus

show 1 reply
luciana1utoday at 8:44 AM

solving ARC-AGI and being useful turned out to be two different problems

show 1 reply
dyauspitrtoday at 7:11 AM

Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.

show 3 replies
rurbantoday at 9:52 AM

Deepseek V4 and Kimi 3 still missing, at least GLM is there.

tonyhart7today at 7:46 AM

cost 20k ???? man

those are like software engineer from third world country

saberiencetoday at 4:06 PM

ARC-AGI is a terrible benchmark for testing LLMs because LLMs are not made, trained, or tuned for playing games.

They are trained on text to respond well to text based questions and do tasks involving modifying text files.

They are not designed for playing games, looking at games, or visual puzzles. Also translating games into text input for the LLM skews the test completely.

Imagine trying to get a human to solve visual puzzle but they can’t look at the puzzle but it has to be explained to them in textual format, we would be terrible at it.

But yet we persist in wasting time on this benchmark. It doesn’t mean anything.

Anoiantoday at 7:14 AM

[dead]

ai_fry_ur_braintoday at 7:27 AM

[dead]