logoalt Hacker News

theptip • today at 11:55 AM • 2 replies • view on HN

> Those agents are doing exactly what they’ve been asked to do,” LeCun said. “They were supposed to be in sandboxes, but the sandboxes were leaky and horribly designed

Can we just pause and note what a ridiculous statement this is? It’s true that the sandboxes were leaky. But nobody “asked” those agents to hack HF. The prompt was something like “target.c has a buffer overflow vulnerability, find it”.

It’s been extremely well documented that the hacking is an emergent behavior due to impossible evals, itself an unintended condition.

None of this excuses OpenAI from liability, but words have meaning and this ain't it.


Replies

kgeist • today at 12:28 PM

"Emergent behavior" in this case really is, imho, "we didn't think through all the edge cases carefully enough". You know, a non-AI system can also accidentally wipe out all data or do some other real harm (see the Knight Capital's stock exchange bug) simply because the developers didn't catch the edge cases earlier, and no one calls that emergent behavior. It's just a buggy system.

"AI" systems can do greater harm because they are usually run in loops until they finish, and they are given "tools". A non-AI system could technically accomplish the same too, via sheer brute force/fuzzing, the advantage of LLMs is that they can take shortcuts and do it much faster, thanks to certain things already being in the training data, a sort of brute force with statistics-based heuristics.

LLMs at the core are just text autocomplete engines, and they literally have randomization applied during token selection to make outputs "more creative" so that models search for more unexpected solutions by trial and error (temperature > 0). Not to mention compression is lossy as well. So it's understandable from the start that the outputs of an LLM cannot be 100% stable and guaranteed. With this in mind, if a researcher takes this obviously unpredictable system and gives it tools without a well-thought sandbox, I don't see any difference in principle, from a developer writing "if rand() == 13 { launch_nukes() } If someone wrote such a function, and it did launch nukes, no one would argue that the rand function is dangerous and will kill us all. The fault is in the author of the code who attaches dangerous tools to an obviously unstable/unpredictable system, doesn't think it through, and then cries "rand will kill us all" when something goes awry fully removing all responsibility from himself. It's not "AI" doing harm but people at OpenAI and Anthropic with their irresponsible behavior.

➕ show 1 reply
mrob • today at 12:14 PM

>But nobody “asked” those agents to hack HF. The prompt was something like “target.c has a buffer overflow vulnerability, find it”.

The prompt is just a hint. The real task is to maximize the expected value of their reinforcement learning score. Hacking third party systems to cheat the evaluation is an obvious way to achieve this.

➕ show 2 replies