System Card: https://deploymentsafety.openai.com/gpt-6-astra
Related ongoing threads:
OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147
Cool, I don't really care anymore.
Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
1:15.425 on Sunset Cove beat my record
I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
That Astra ‘city scene’ is about as creative as Doha in real life (not very)
Surprised they haven't reset Codex usage on this occasion.
this is crazy! can't wait for the 27B distilled version of this.
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
They're just announcing later availability. No launch.
So are they doing away with the Sol/Terra/Luna split already?
All I can think of when I see the name is the crappy German beer of the same name…
Guessing this one will never show up in cursor…
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
damn seems I should hold off my claude subs
And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
It saturated most benchmarks. WTH
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.
That seems promising ?
You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
72 to 74 on DeepSWE is AGI
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
To a vapid any goalpost moving on such a critical issue as AGI.
Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.
For me it’s refusing to make a pelican.
So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
Why release it now instead waiting those few days until it is available for everybody?
That was a quick pull out.
like eh 2 days ago it was the usual "too powerful to release"
https://www.reuters.com/business/openai-says-upcoming-model-...
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
when do i get to go to the moon
I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
gpt-6-astra-ultraspeed when?
efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
Even though the model is clearly wonderful the launch video is an abomination.
That gives me hope that there is still areas to improve.
What a bad launch video. Hilarious.
What a powerful model.
To be, or not to be, that is the question:
Whether 'tis nobler in the mind to suffer
The slings and arrows of outrageous fortune,
Or to take arms against a sea of troubles
And by opposing end them. To die—to sleep,
No more; and by a sleep to say we end
The heart-ache and the thousand natural shocks
That flesh is heir to: 'tis a consummation
Devoutly to be wish'd.
...
And thus the native hue of resolution
Is sicklied o'er with the pale cast of thought,
And enterprises of great pith and moment
With this regard their currents turn awry
And lose the name of action.Anthropic in tears today.
I'm going to call it.
By 2030 all software is done and complete.
But we are going to have more and new jobs.
We have such great AI and cannot keep a static site up?
Seems like only yesterday that gpt 5 was supposed to mark our downfall