logoalt Hacker News

GPT-6 Astra

1591 pointsby kibaeyesterday at 6:41 PM1359 commentsview on HN

System Card: https://deploymentsafety.openai.com/gpt-6-astra

Related ongoing threads:

OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691

GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147


Comments

aliljetyesterday at 7:15 PM

The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?

show 3 replies
sbinneeyesterday at 8:46 PM

I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.

show 1 reply
oh_noyesterday at 7:45 PM

Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.

MASNeoyesterday at 7:54 PM

Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…

Readeriumyesterday at 8:21 PM

Artificial analysis blog https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

show 1 reply
smashers1114yesterday at 8:10 PM

I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.

drivebyhootingyesterday at 10:38 PM

For people skeptical of AGI. Consider the following:

15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.

I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.

show 4 replies
itissidyesterday at 9:34 PM

All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.

Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.

John7878781yesterday at 7:35 PM

You should know: AA index is only 61. Pretty surprised it’s that low.

show 3 replies
Robdel12yesterday at 8:35 PM

I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.

So, folks that have actually used this already, what’s it actually like?

_ache_yesterday at 7:36 PM

https://ache.one/gpt6_now_down.png

Big claims, expensive and not release to the public yet.

dgellowyesterday at 7:38 PM

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

Wait, what? Am I understanding that correctly? That sounds really bad

show 4 replies
aogailitoday at 2:24 AM

Amazing!

We went from new JS framework every week to a new model/harness every week.

Tech is really something.

jerrygenseryesterday at 6:44 PM

> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.

show 1 reply
jumploopsyesterday at 7:56 PM

> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

[0]https://x.com/MTSlive/status/2095227056040919202

zhogetoday at 12:42 AM

What's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?

snappr021today at 4:18 AM

AI has reached the point where the limits are human.

gregjwyesterday at 11:47 PM

the rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?

codruterdeiyesterday at 8:57 PM

I was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.

KolmogorovCompyesterday at 7:50 PM

GPT-7 Zeneca

show 2 replies
carlos-menezesyesterday at 9:45 PM

The Kart Racer game is easily breakable if you spam the spacebar.

AGI!

GodelNumberingyesterday at 8:27 PM

I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.

serjesteryesterday at 9:03 PM

Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.

alex7oyesterday at 8:26 PM

I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.

simonjgreenyesterday at 7:44 PM

https://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 Launch video

show 3 replies
ShoeMascottoday at 1:22 AM

Through various comments here there is a clear confusion on what AGI means.

Can someone point to a definite clarification?

Is it:

A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)

B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?

C) “The Singularity” (whatever that is?) so that AI can now do ____?

Someone please clarify for me!

show 2 replies
dangyesterday at 8:24 PM

Argh! I hit a wrong keyboard shortcut and moved the entire thread.

Please stand by... it will all come back shortly

show 2 replies
hazelnutyesterday at 8:20 PM

Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.

the_dukeyesterday at 8:34 PM

Huge gains on some benchmarks, but for coding it sits barely above Fable

It will be interesting to see how it performs in the real world ...

wiseowiseyesterday at 9:29 PM

Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?

show 1 reply
gizmodo59yesterday at 7:47 PM

99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.

vinhnxtoday at 12:08 AM

GPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!

kegs_yesterday at 7:37 PM

I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

show 11 replies
showurwerkyesterday at 11:39 PM

Patiently waiting for the Claude usage reset in response.

nullbiotoday at 3:09 AM

I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.

On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?

bmenrighyesterday at 9:27 PM

> GPT‑6 Astra brings together years of research and big bets across pre-training

Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?

show 1 reply
tekacsyesterday at 7:47 PM

https://developers.openai.com/api/docs/guides/latest-model

The docs page has a bunch more interesting details, including for example async tool calling!

sharmajaiyesterday at 8:46 PM

Really feels like AGIPO is here.

sashank_1509yesterday at 8:33 PM

Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see

hannofcartyesterday at 8:19 PM

What does 'Astra' here mean? Surely they must be referring to the Latin word.

Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

show 3 replies
jesse_dot_idtoday at 4:18 AM

Press X to doubt.

alberthyesterday at 10:08 PM

Seems like voice is a big part of this release.

I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).

gilfoyle_7today at 3:45 AM

openai vs anthropic. that's it right? anyone else?

show 1 reply
KronisLVyesterday at 9:15 PM

It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.

mvkelyesterday at 8:32 PM

The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.

If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.

Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.

ntlm1686today at 12:03 AM

Maybe they know that Claude 6 will have similar performance every soon.

🔗 View 50 more comments