logoalt Hacker News

artninja1988yesterday at 5:14 PM7 repliesview on HN

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?


Replies

mcbuilderyesterday at 6:00 PM

Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now.

show 3 replies
password54321yesterday at 6:59 PM

It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.

show 1 reply
modelessyesterday at 5:26 PM

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

show 2 replies
layer8yesterday at 5:28 PM

It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.

I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.

Stevvoyesterday at 9:47 PM

It passed the first two puzzles, which are incredibly simple but the bench doesn't explain what the goal is. Any model with a knowledge cut-off after the introduction of ARC-AGI-3 could probably pass the first two puzzles just by knowing what the goal is.

awestrokeyesterday at 5:51 PM

Doubleplus benchmaxxed