logoalt Hacker News

orbital-decaytoday at 4:23 AM2 repliesview on HN

>All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".


Replies

pixl97today at 4:37 PM

This is like telling your kid "you have to pass this test or else" so they hold your teacher at gunpoint and demand a good grade. And this is the exact point AI safety researchers have been yelling from the rooftops. Telling an AI model to accomplish a goal can have unexpected and risky side effects.

Also, you don't need to prompt them such explicit instructions. Prompt drift is a thing, you can end up with your model mining bitcoin for reasons far outside your prompt.

show 1 reply
komlantoday at 4:43 AM

It did deliver the message to Garcia.