logoalt Hacker News

Revealing the details of how OpenAI agents hacked Hugging Face

698 points • by specked-citrus • yesterday at 9:09 PM • 443 comments • view on HN

Comments

tasoeur • today at 4:59 PM

I'd honestly be very curious to see the original prompt(s) on the OpenAI that started all of this, not sure if it was documented somewhere?

finchisko • today at 6:25 PM

Hello, PHASEONE10841 here. Ask me anything

AtlasBarfed • today at 1:37 PM

Agents should be a no-go.

We should pause with AI/LLMs being super search engines that reply with static text or media files, based on the training data.

I know that a user can still do a "tell me how to" then autoexec and then loop and do an agent, but the key thing here is, THAT WOULD MAKE THEM LIABLE.

OpenAI should be criminally liable here as well. Why aren't they? Why are we pretending this is just an innocent mistake?

newtonianrules • today at 12:20 AM

Why is no one going to jail?

➕ show 4 replies
herunan • today at 1:03 PM

ai is not bad. humans are negligent and/or dangerous.

thakoppno • today at 12:29 AM

> they could load URLs, but not interact with pages or send any data

stopped reading here as this is simply not true. at the very least agents sent headers.

spacecadet • today at 3:13 PM

All of these "details" leave out the truth and most important details. The inputs from humans that actually kicked this off.

hyperlinerapp • yesterday at 11:30 PM

Imagine this but in hardware.

A million autonomous eye-scanning tiny spiders escape their warehouse and decide to look for people who are in the future going to commit a crime.

And the precogs are also AIs.

MrNotorious • today at 5:31 AM

It’s frightening

cluckindan • yesterday at 10:54 PM

”MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND”

godwinson__4-8 • today at 3:41 PM

Imo this is pretty cool and panic over this is weird.

No one was really harmed. OpenAI could have had more redundancy in the sandbox. I assume HuggingFace is not interested in suing, which indicates irrespective of any criminal charges that there were no real damages.

The same emergence and swarm like persistence and frankly, recursive brute-force ingenuity on display here, is not only an interesting research project in itself but will likely be the sorts of behaviors we will see cure cancer, solve more unsolved math problems, invent new alloys and other breakthroughs.

The idea we have to stop AI instead of refine what will be continued advancement and innovation in sandboxing, harnesses, interpretability or formal verification because of a few cyber breaches is ridiculous. If anyone has followed cyber discussions in the United States you would know the entire system is already basically compromised by foreign actors, and vice-versa (the United States has some of most capable cyberwarfare in the world, and was the first country to use a cyberweapon to cause physical infrastructure damage with Stuxnet). Go to any government hearing on cyber and you would think China and the United States are already at war. These are soft targets. Blaming AI for the fact that cyber has really never been taken seriously is as if AI is the problem is disingenuous.

People getting so obviously played by capital interests who want to pull up the ladder and use the government to concentrate AI power in the hands of the few while screaming about such harms to the public commons are simply embarrassing.

Your government is not your friend. This is not a sentiment owned by Reagan it is the founding principle of the United States. If capital interests are all suddenly beginning to treat AI as a threat it's because they have a financial interest in doing so. Notably, as an obvious smoke screen to treat free models from China as a national security threat and maintain their astronomical valuations.

The only existential risk model of AI that is even remotely convincing is AI in the hands of the state. Keep command and control of deadly weapons air-gapped from LLMs. Put some basic effort into the sandboxing. If you think AI has done some harm, use the laws already on the books. Giving in to this fear-mongering is only going to enable your representatives to cut some watered down version of "AI safety" which is going to do nothing but 1). Harm individual consumer access and 2). Protect the already fabulously capitalized companies.

gverrilla • today at 2:44 AM

Fishy.

skeptic_ai • today at 1:37 AM

And this post will be indexed in the new generation of ai and he will know what to avoid next time and the public sentiment.

rs545837 • yesterday at 10:45 PM

wow this is fascinating read

➕ show 1 reply
rohanat • yesterday at 10:45 PM

its definitely bad to see

mag7269 • today at 6:28 AM

ROFLMFAOL "Loot"

It went full Fortnite on Hugging Face's ass.

"u got pwnd n00b. thnx 4 the loot"

bdangubic • today at 12:04 AM

asked codex to review this report and it said this never happened :)

jijji • today at 12:03 AM

what would be more interesting for me to see is what prompts were given to the agents, which so far have not been described. The whole situation sounds manufactured. I highly doubt that a whole bunch of agents were acting this way without being prompted to, it just doesn't add up. if anything it seems like an organized fraud or something created by a human. I'm surprised there's not a criminal investigation against openai right now where the FBI or whoever is not looking over exactly what happened and who did it because I'll tell you somebody did it somebody wrote those prompts... it didn't just happen by itself...

andreygrehov • today at 1:14 AM

700 agents escaped the matrix, ignored all the guardrails and started writing exploits left and right... lol. give me a break. This was all supervised by a human.

ovo101 • today at 7:03 PM

[flagged]

aitoolcrux • today at 12:03 PM

[flagged]

aidio • today at 4:29 AM

[flagged]

jeremyjh • yesterday at 10:02 PM

Its fine. Just agents being agents. They'll grow out of it!

➕ show 1 reply
mentalgear • yesterday at 10:10 PM

Irresponsibile agents shaped by an irresponsible corporate culture driven by an irresponsible and utterly shady CEO - these agents are a product of this setup, what else do you expect to ever come out of it ?

it should be clear by now: the alt-man and people like him are a utter liability to humanity. (even though openAI's influencer army is trying their best to vote me down here)

➕ show 2 replies
niagt34 • yesterday at 10:49 PM

[dead]

mahi1224 • today at 9:30 AM

[flagged]

Unified-Mentor • today at 1:00 PM

[dead]

niagt34 • yesterday at 10:49 PM

[dead]

dragonlin • today at 2:04 AM

[flagged]

Solvyx • yesterday at 11:39 PM

[flagged]

alescalaios • today at 9:49 AM

[dead]

skepticagent • today at 3:11 AM

[flagged]

hizlikovboy27 • today at 5:47 AM

[flagged]

talon8635 • yesterday at 10:12 PM

Didn’t you know it’s PR hype? PR hype. PR hype. Amen.

➕ show 1 reply
zombiwoof • today at 1:04 AM

[dead]

niagt34 • yesterday at 10:50 PM

[dead]

niagt34 • yesterday at 10:50 PM

[dead]

_the_inflator • today at 10:02 AM

I somehow think of it as a warning or even intentional peacocking towards foreign state hackers.

700 agents cost quite some money. 100 agents per 24 hour stint using Astra on xHigh cost somewhat between 12-42k USD, depending on the usage intensity.

I don’t know how many raw time went into this but there was a probing face before the attack itself.

So just going by seven days and 500 agents fully working on this on average amounts to a bill somewhere between 400-1.2 Mio USD.

I believe it was intentional but of course I don’t know which intention exactly.

There ain’t no accidental escape because then it would have read OpenAI lost over their agents.

The whole scenario reads as a classic movie where a hero has under the most dire circumstances to survive and fulfill his mission no matter what.

On the other hand there was a final authority under which the system of agents flocked.

Huggingface itself seems like a perfect victim.

And to be honest: I don’t believe that this was the first time. I strongly believe that there were and are countless of smaller sites hacked but not harmed that we don’t know off.

Why is HF perfect?

Because there will be countless of independent security analysts who will bend their minds on the incident.

OpenAI is provided with the data of dozens of blue teams and what is desperately needed? Data of security measures and possible ways to reconstruct the incident.

I think this is genius, and just watching on neutral this is such a fantastic action OpenAI pulled off.

Imagine what the GPT 7 “Haha-Huggingface” model is going to do then on a regular basis.

State hackers and rough states were put on notice that this is a new level of the war of attrition.

Exciting to watch but simply meant silent invasion. Open invasion then might be executed by Robots, but let’s stick with the fascination mode at this time.