logoalt Hacker News

tintortoday at 2:12 AM3 repliesview on HN

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/

- Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra

Who is wrong here?

Some benchmark results in Astra page for Fable and Opus are blank (-).

What is Artificial Analysis intelligence index measuring that Astra scores poorly on?

Can someone from OpenAI / Artificial Analysis comment / clarify?

Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.


Replies

dannywtoday at 3:18 AM

I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.

That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.

But AA scores Gemini 3.8 Flash at 59, and Astra at 61.

kubrickslairtoday at 2:19 AM

Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.

Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.

show 1 reply
AnodicElegytoday at 2:40 AM

If you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.