logoalt Hacker News

stratos123yesterday at 6:57 PM0 repliesview on HN

I'm not even sure they are. This incident isn't that much different from the OpenAI swarm Huggingface hack incident - and in that one, all the models involved (despite being internal) were safety-trained. It seems what the safety training amounts to is (as the METR report puts it) "expressing ethical hesitation" before going along with it anyway.