"Crimes" or even "hacking" are not really that impressive. I can get GPT-2 to abet financial fraud or write exploits with the right prompt. Some people get accused of hacking crimes for just using Inspect Element. It's a moving goalpost with some very low bars to cross.
OpenAI's adversarial agent was caught almost immediately, and the entire thing was rushed out as a press release. It reads like a clickbait lab experiment more than an actual alignment concern.
> I mean the model committed numerous crimes
"Crimes" or even "hacking" are not really that impressive. I can get GPT-2 to abet financial fraud or write exploits with the right prompt. Some people get accused of hacking crimes for just using Inspect Element. It's a moving goalpost with some very low bars to cross.
OpenAI's adversarial agent was caught almost immediately, and the entire thing was rushed out as a press release. It reads like a clickbait lab experiment more than an actual alignment concern.