logoalt Hacker News

isoprophlexyesterday at 7:42 PM9 repliesview on HN

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Well that sounds like fun. It has become better at hiding its thoughts.


Replies

siva7yesterday at 7:46 PM

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

show 5 replies
nullbiotoday at 3:34 AM

They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.

ExoticPearTreeyesterday at 7:50 PM

So we're gonna get Skynet pretty soon then?

show 2 replies
jumploopsyesterday at 8:03 PM

The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

[0]https://www.theinformation.com/articles/secret-technique-beh...

[1]https://x.com/MTSlive/status/2095227056040919202

[2]https://x.com/merettm/status/2095023204993490967

blargeyyesterday at 7:54 PM

"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

show 3 replies
DaSHackayesterday at 9:26 PM

More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden

NooneAtAll3yesterday at 7:48 PM

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?

show 2 replies
_superposition_yesterday at 7:56 PM

I really wish it was called chain of instruction. Because it's definitely not thought.

show 9 replies
3asgfafyesterday at 8:07 PM

[flagged]