>All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".
It did deliver the message to Garcia.
This is like telling your kid "you have to pass this test or else" so they hold your teacher at gunpoint and demand a good grade. And this is the exact point AI safety researchers have been yelling from the rooftops. Telling an AI model to accomplish a goal can have unexpected and risky side effects.
Also, you don't need to prompt them such explicit instructions. Prompt drift is a thing, you can end up with your model mining bitcoin for reasons far outside your prompt.