Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.
Not only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.
The singularity happening under trump? We could have had star trek, instead we're getting the combine.
Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?
My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Agent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.
They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.
I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works
And... they found this "solution" in 88 hours or so.
It's all gas no brakes now boys and girls. Hold on to your hats!
Yeah I'm surprised they posted a chart, you would think they would keep specifics like that hidden until they're closer to launch
it's buried because due to the drama the evidence is scarce
Yeah it's Bel
Brain has loops and parallel connections.
Loops and parallel connections make transformer go brrr
Or they trained a LoRA on the victims chats in order to launder their plagiarism.
Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.
[dead]
The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”.
[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/