logoalt Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

186 pointsby vignesh_wararyesterday at 7:45 PM118 commentsview on HN

Comments

at1astoday at 12:04 AM

I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken.

From https://epoch.ai/latest/announcing-frontiermath-erdos

> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours

> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.

Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.

malfistyesterday at 8:37 PM

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

show 8 replies
Betelbuddyyesterday at 8:03 PM

"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

Well I dont know about all of you, but I am celebrating meat based humans...

show 2 replies
modelessyesterday at 11:33 PM

$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.

fastballyesterday at 10:31 PM

Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?

an0malousyesterday at 11:22 PM

Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.

show 1 reply
6thbityesterday at 9:32 PM

The instant/no reasoning performed extremely well

    none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.
show 1 reply
dwohnitmokyesterday at 8:28 PM

> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

mikert89yesterday at 8:48 PM

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

show 3 replies
fxdyesterday at 10:40 PM

“AGI” never made sense to me. It’s a purely marketing term right?

I’ve ignored it thinking it would go away, but it keeps coming up.

I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.

Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.

But you’d be no nearer to solving consciousness.

Given this thought trajectory - what is AGI supposed to be?

show 5 replies
scotty79today at 3:41 AM

How good are LLMs at doing Mensa tests?

show 1 reply
piloto_ciegoyesterday at 8:17 PM

99.9% with the right harness? Ok, we're at AGI then.

Prediction:

We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.

show 6 replies
yomismoaquiyesterday at 10:33 PM

Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI?

Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

hypferyesterday at 8:49 PM

What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?

show 1 reply
yusufozkanyesterday at 7:49 PM

what the hell is that score/cost curve lol

show 1 reply
ajjahsyesterday at 11:31 PM

[dead]

bigbuppoyesterday at 9:46 PM

Wake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.