logoalt Hacker News

simonwyesterday at 5:30 PM1 replyview on HN

They didn't have watchdog agents - those exist for their production models but had been deliberately removed for the purpose of this evaluation.

OpenAI wrote about how their mechanism for that in production works here: https://openai.com/index/safety-alignment-long-horizon-model...

> We created a monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary. The monitor observes not just a single action but the entire trajectory.


Replies

avadodinyesterday at 11:10 PM

So core ChatGPT models are all abliterated?

Open weights releases should have a companion model that censors boobs and Tiananmen and then we would have a useful model and something we could fine–tune into a useful model instead of a half–useful model needing a lobotomy.

And I mention Tiananmen because I feel like the Western models are built easier to abliterate judging by the quality difference.