They didn't have watchdog agents - those exist for their production models but had been deliberately removed for the purpose of this evaluation.
OpenAI wrote about how their mechanism for that in production works here: https://openai.com/index/safety-alignment-long-horizon-model...
> We created a monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary. The monitor observes not just a single action but the entire trajectory.
So core ChatGPT models are all abliterated?
Open weights releases should have a companion model that censors boobs and Tiananmen and then we would have a useful model and something we could fine–tune into a useful model instead of a half–useful model needing a lobotomy.
And I mention Tiananmen because I feel like the Western models are built easier to abliterate judging by the quality difference.