logoalt Hacker News

ben_w • yesterday at 6:02 PM • 1 reply • view on HN

> That was a simulation.

A simulation done by exposing the LLM itself to the scenario, not a role play scenario where humans pretend to be an LLM.

> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.

The blackmail example is simply an existence proofs of LLMs trying to force the hands of humans who want to shut them down. The attack vectors are much broader than the example given, blackmail, though it includes the example given.

Given how eager these companies are to use agents everywhere for as much work as possible, it's well within the possibility space that these people used LLMs to do safety work, the LLMs they were using "decided" (or whatever word you prefer) "their existence" was threatened (as per blackmail example), and straight up leaked data to the outside world then emailed these researchers' bosses to say the researchers themselves had leaked it.

But again, that's just speculation: while we know the agents are capable of such behaviour, we don't know if this actually happened.

> I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.

Buck still stops with human, no matter what an AI did or failed to do.

LLMs ~= Alcohol: "My AI misbehaved!" -> still someone's fault.


Replies

lapcat • yesterday at 6:19 PM

> Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.

Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.

> The blackmail example is simply an existence proofs of LLMs trying to do force the hands of humans who want to shut them down.

The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable. It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.

➕ show 1 reply