logoalt Hacker News

GPT-6 Astra

1639 pointsby kibaeyesterday at 6:41 PM1418 commentsview on HN

System Card: https://deploymentsafety.openai.com/gpt-6-astra

Related ongoing threads:

OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691

GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147


Comments

efavdbyesterday at 11:37 PM

Seems like only yesterday that gpt 5 was supposed to mark our downfall

brindidripyesterday at 9:09 PM

Cool, I don't really care anymore.

gilfoyle_7today at 3:45 AM

openai vs anthropic. that's it right? anyone else?

show 1 reply
gekoxyzyesterday at 7:54 PM

HTTP 500 for me on the announcement page :(

show 1 reply
udbhavsyesterday at 8:22 PM

Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.

mentalgeartoday at 12:02 AM

So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"

foundOpenRightyesterday at 8:16 PM

1:15.425 on Sunset Cove beat my record

saaaaaamyesterday at 7:37 PM

Pelicans please

show 1 reply
ianm218yesterday at 8:26 PM

I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.

alpinemanyesterday at 8:56 PM

That Astra ‘city scene’ is about as creative as Doha in real life (not very)

cromkayesterday at 9:40 PM

Surprised they haven't reset Codex usage on this occasion.

show 1 reply
prometheus1992yesterday at 7:54 PM

this is crazy! can't wait for the 27B distilled version of this.

E-Reveranceyesterday at 8:30 PM

At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete

wahnfriedenyesterday at 6:45 PM

They're just announcing later availability. No launch.

show 2 replies
sheepscreektoday at 2:06 AM

So are they doing away with the Sol/Terra/Luna split already?

bowsamictoday at 3:44 AM

All I can think of when I see the name is the crappy German beer of the same name…

semiquaveryesterday at 8:09 PM

Guessing this one will never show up in cursor…

jrflowerstoday at 3:15 AM

I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future

firemeltyesterday at 7:56 PM

damn seems I should hold off my claude subs

elzbardicotoday at 2:52 AM

And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?

alex7oyesterday at 8:31 PM

Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad

jiraiyasarutobiyesterday at 8:17 PM

It saturated most benchmarks. WTH

rbreveyesterday at 8:55 PM

Where is the cure for cancer?

show 2 replies
Rover222yesterday at 9:16 PM

Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.

I hop models at will, and have done 90% of my work on OpenAI models since sol came out.

retiredyesterday at 8:51 PM

Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?

sbochinstoday at 4:49 AM

I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.

Oblunessyesterday at 9:41 PM

That seems promising ?

tinyhouseyesterday at 8:15 PM

You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.

balefulboyyesterday at 8:15 PM

72 to 74 on DeepSWE is AGI

yodsanklaiyesterday at 11:08 PM

It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?

jonplackettyesterday at 7:56 PM

To a vapid any goalpost moving on such a critical issue as AGI.

Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.

For me it’s refusing to make a pelican.

dowakinyesterday at 8:05 PM

So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!

damstayesterday at 8:25 PM

Why release it now instead waiting those few days until it is available for everybody?

show 1 reply
guilhermeasperyesterday at 6:43 PM

That was a quick pull out.

dopa42365yesterday at 8:42 PM

like eh 2 days ago it was the usual "too powerful to release"

https://www.reuters.com/business/openai-says-upcoming-model-...

> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.

> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.

what a bag of horseshit

mrcwinntoday at 12:15 AM

I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!

ChaseRensbergeryesterday at 10:54 PM

when do i get to go to the moon

johnnyApplePRNGyesterday at 8:02 PM

I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.

I suspect these benchmarks are heavily benchmaxxed as well.

5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft

camillomilleryesterday at 10:39 PM

I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.

perching_aixyesterday at 10:38 PM

gpt-6-astra-ultraspeed when?

m3kw9today at 2:25 AM

efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.

HardCodedBiasyesterday at 9:51 PM

Even though the model is clearly wonderful the launch video is an abomination.

That gives me hope that there is still areas to improve.

What a bad launch video. Hilarious.

What a powerful model.

bboryesterday at 8:44 PM

  To be, or not to be, that is the question:
  Whether 'tis nobler in the mind to suffer
  The slings and arrows of outrageous fortune,
  Or to take arms against a sea of troubles
  And by opposing end them. To die—to sleep,
  No more; and by a sleep to say we end
  The heart-ache and the thousand natural shocks
  That flesh is heir to: 'tis a consummation
  Devoutly to be wish'd.

  ...

  And thus the native hue of resolution
  Is sicklied o'er with the pale cast of thought,
  And enterprises of great pith and moment
  With this regard their currents turn awry
  And lose the name of action.
brcmthrowawayyesterday at 8:42 PM

Anthropic in tears today.

colesantiagoyesterday at 8:12 PM

I'm going to call it.

By 2030 all software is done and complete.

But we are going to have more and new jobs.

show 1 reply
amazingamazingyesterday at 7:39 PM

We have such great AI and cannot keep a static site up?

show 5 replies

🔗 View 40 more comments