logoalt Hacker News

HarHarVeryFunnyyesterday at 10:00 PM1 replyview on HN

> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo

There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.

https://alignment.openai.com/measuring-reward-seeking/

Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.

Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.


Replies

camgunzyesterday at 11:49 PM

Yeah. It's a lot easier to destroy than create, and though I think LLMs are mostly shit at creating, they're much better at the simpler destroy task. To be clear, we don't know and probably can't know everything that happened with this incident. We unleashed thousands of highly capable, autonomous, unpredictable, well-resourced programs onto the open internet for an extended period of time. We are in no way treating this with the seriousness it deserves, because the stock market essentially depends on this garbage and the current US is miserably incompetent.