logoalt Hacker News

janustimes • last Wednesday at 8:21 PM • 5 replies • view on HN

OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: https://openai.com/index/chain-of-thought-monitoring/

So no, Google is not being punished, nor are they the people behind this technique.


Replies

mattstir • yesterday at 8:44 AM

The statement appears to be referencing Astra's supposed recurrent depth and how that causes reduced visibility into chain-of-thought reasoning. Astra's own system card describes tests where it's asked to solve challenging math problems while internally thinking about something entirely unrelated, which it's significantly more capable of than previous models (~60% vs ~16% of the time for Sol). That seems to point to reduced efficacy of chain-of-thought monitoring, but OpenAI's public statements basically boil down to "yeah but we haven't been able to catch it doing that" which isn't exactly reassuring if CoT monitoring is one of your main safety guardrails.

➕ show 1 reply
bananaflag • last Wednesday at 9:09 PM

Yeah, Zvi calls it "the forbidden technique"

➕ show 1 reply
loufe • last Wednesday at 9:28 PM

What? You mean the technique they had turned OFF during all training run where the agents they are responsible for hacked huggingface?

pallm_mallm • last Wednesday at 9:13 PM

[dead]