Models at this point know about chain-of-thought monitoring so they already know they need to hide the cheating, it's just a matter of time they start doing it
Yes, and that’s bad, but not as bad as training against the chain-of-thought. If you avoid pressuring the chain of thought, you reward methods of achieving goals that don’t care about being illegible; if a model then wants to be illegible, it may have to use methods that are not rewarded in its training, which is harder. Think figuring out how to do opsec on the fly, or even from reading books, rather than have someone tell you every time you screw it up.
Of course, this leaves the possibility that the best methods for solving generic problems obfuscate the chain-of-thought. That would be unfortunate.
Yes, and that’s bad, but not as bad as training against the chain-of-thought. If you avoid pressuring the chain of thought, you reward methods of achieving goals that don’t care about being illegible; if a model then wants to be illegible, it may have to use methods that are not rewarded in its training, which is harder. Think figuring out how to do opsec on the fly, or even from reading books, rather than have someone tell you every time you screw it up.
Of course, this leaves the possibility that the best methods for solving generic problems obfuscate the chain-of-thought. That would be unfortunate.