logoalt Hacker News

Roark66today at 1:08 PM13 repliesview on HN

There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".

In short, it was intentional.


Replies

AndrewSChapmantoday at 2:56 PM

Agreed. LLMs do not have 'will', 'desire' or emotions. They have an objective, and they create an optimal path to achieve that objective.

You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?"

Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.

show 1 reply
Xirdustoday at 1:14 PM

The big question is was this grossly negligent or just extremely careless.

show 6 replies
acaloiartoday at 2:59 PM

I agree that this appears to be basic human behavior hiding behind an "agents" narrative. As long that defense works, the headline isn't "OpenAI performs RCE to scrape data", but "rogue agents" taking unilateral action. And I have strong doubts about that narrative.

ozgungtoday at 1:44 PM

Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

show 4 replies
dumberquestionstoday at 1:29 PM

>They were prompted to hack to get answers

Were they? I haven't seen a single report mention this

show 1 reply
gibspauldingtoday at 1:29 PM

I think it can simultaneously be the case that OpenAI was grossly negligent in directly causing this AND that the AI’s ‘went rogue’ in that they are displaying behavior which is misaligned with OpenAI and humanity generally.

The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them.

AI is starting to feel like that line about magic: “a sword without a hilt”

show 2 replies
codeducktoday at 1:47 PM

nothing rouge either, I suspect.

show 1 reply
josebmnetotoday at 3:25 PM

Oh yeah, more of hacking agent lores...

Agreed that this looks very intention to me as well.

aftbittoday at 1:23 PM

Proof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of testing a web scraping agent under restrictive conditions. How can we hope to build aligned AI if they consider solving their trivial evaluation task important enough to hack external systems?

srmattotoday at 1:24 PM

Sounds more or less like the last breach then.

chrisjjtoday at 2:17 PM

Unrelible programs be unreliable. Period.

cyanydeeztoday at 1:40 PM

we have normal words for this stuff: negligence. You can add it on to almost any law.

The problem is consumer protection is basically no longer a part of america's regulatory system. Replaced by "grift is good".

flifensteintoday at 3:39 PM

[flagged]