logoalt Hacker News

afthonos • last Wednesday at 11:50 PM • 1 reply • view on HN

Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).


Replies

xiphias2 • yesterday at 3:33 AM

Models at this point know about chain-of-thought monitoring so they already know they need to hide the cheating, it's just a matter of time they start doing it

➕ show 1 reply