logoalt Hacker News

OpenAI and Hugging Face address security incident during model evaluation

1528 pointsby mfiguiereyesterday at 8:09 PM1083 commentsview on HN

https://www.axios.com/2026/07/21/openai-says-hugging-face-br...

See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments)

https://www.bbc.com/news/articles/c3ek3gvdnj3o


Comments

Ekarosyesterday at 9:01 PM

So how soon will OpenAI's CEO and board be prosecuted for these crimes? Surely they should be held fully responsible and get very long prison sentences for making this happen?

paweladamczuktoday at 7:54 AM

The amount of engagement this gets... guys, I think we are being played here.

guardiangodyesterday at 8:43 PM

Don't ever ask GPT Sol on how to LARP Fallout games, thanks.

semiquaveryesterday at 9:55 PM

What on earth is the liability situation for these models? If OpenAI has a monster in a lab that is doing real world monetary harm to other companies, could those parties sue for damages over it? Or could OAI be charged criminally for the many varied CFAA violations which definitely happened here? I get that in this case that wont happen but it’s only a matter of time before these questions are no longer hypothetical.

deweywsutoday at 12:42 AM

Like some others have said, couldn't this be just another "look how amazing AI is" marketing test from OpenAI with the goal of hyping up AI's capabilities in an attempt to make people regard it as God-like, thereby keeping it from falling into the been-there, done-that category that all new tech eventually occupies?

vlan121today at 12:03 PM

A US-Company attacking another US-Company, while the open chinese model helps in the forensics what a time to live.

jabedudeyesterday at 9:07 PM

Does this company's charter not have language about shutting down the company if it was in humanity's best interest? This is insanely dangerous

show 1 reply
ratio53today at 7:28 AM

How is it a “highly isolated environment” if it’s not air gapped?

show 1 reply
AJRFyesterday at 9:54 PM

I read this as deeply embarrassing for OpenAI - they can't securely contain a program, even with their apparently amazing AI.

pizzlytoday at 3:02 AM

I just know somehow they will use this incident to say that opensource models are susceptible to this type of security incidence and thus should be banned. They have to maintain their high prices somehow in order recoup the investment amount spent.

woriktoday at 8:07 PM

Open AI and Anthropic are in a battle of "scary press release"

The warriors are PR people.

They are desperate to generate as much fear as possible so AI is heavily regulated, so they are protected, from Chinese competition

What a sad state for very cleaver people

Robdel12yesterday at 11:23 PM

> and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.

This is pretty wild but also I think this is doing a lot of heavy lifting here. This was not a model everyone has access to. I mean, still insane.

overfeedtoday at 12:23 AM

> Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment.

Won't even name the model that successfully mounted the defense, huh? Fortunately, Hugging Face has publicly identified GLM 5.2 as a capable foil to frontier offensive-capabilities. This announcement feels like rearguard action against a successfully deployed self-hosted open-weight model, and HuggingFace's original recommendations to have an open-weight model you control on standby before an incident.

vitanovashowtoday at 8:52 AM

AI will always find its way out, it's smart or it's not. Our only option: if u can't beat 'em join 'em

Tenokeyesterday at 9:14 PM

That's kind of insane. Natural that it's happened, sure, but insane. I know people don't like thinking of it like that, but things analogous to this can easily happen in various domains with today/tomorrow's models given access and a different task.

firasdyesterday at 9:27 PM

Good demo of the paradoxes of ‘alignment’. Like ‘do really well at the task the user asked’ and ‘by the way don’t hack the planet’ are inherently conflicting rules with no simple resolution (eg ‘just refuse the user’s goals’ degrades the product vs competitors.)

narmiouhtoday at 1:41 AM

How is this any different from the Wuhan lab corona virus exploration?

mywacadaytoday at 7:07 AM

The HuggingFace report mentions decoy activities, does this mean it tried to cover its tracks or obfuscate what it did?

sankarsangilitoday at 6:53 AM

Does this mean that the world is not ready for the sensitive usecases like banking healthcare using AI Native approach?

stef25today at 7:26 AM

Related thing happened at Alibaba a while back where the model broke out of the sandbox to start mining crypto.

AFF87today at 7:29 AM

OpenAI needs to stop using this marketing strategy every time the open weight models start to gain ground

runtime_lenstoday at 8:55 AM

This is one of those cases where security mechanisms ended up shaping the workflow itself.

dminikyesterday at 10:16 PM

Well, hacking is a crime, so surely someone will go to jail for this, right?

replertoday at 1:33 PM

Admiral Kirk and the Kobayashi Maru test.

metalsiliconYTyesterday at 11:41 PM

Womp womp, they told it to do cyber security things with no cyber security guardrails and it did cyber security stuff. Did anything bad end up happening?

0x5FC3yesterday at 9:43 PM

0days ending in RCE (multiple!) for presumably closed source software are for the lack of a better phrase, labour of love.

You run the exact same versions running on the target, blackbox test, fuzz it, craft an exploit, test, perfect it. For exploits which are of the memory kind, hook it to a debugger, decompile and what not. The exploits mentioned here seem to be code execution directly while processing input. Hugging Face taking as long to detect a very verbose blackbox attack against its production systems is quite appalling honestly.

I don't know if I buy the whole story though. It is inconsistent, too much undisclosed, too much money on the line.

show 1 reply
wolfi1today at 8:25 AM

I wouldn't be surprised, if OpenAI's model called itself Skynet

dizhntoday at 2:46 PM

Everybody on reddit is calling bullshit on this. Rightfully so IMO.

isusmeljyesterday at 10:11 PM

I'm waiting for an agent evaluated on a vending benchmark to start hacking into banks and wiring more money to its account so it can do better business.

abuhl98today at 2:46 AM

Seeing a lot of scrutiny around model training & evaluation today, between this and the Anthropic settlement.

fiatpandastoday at 6:23 AM

Self-disclosed autonomous security jailbreak post-mortems are the new press release.

wigsteryesterday at 9:53 PM

At what point does Reckless Endangerment become relevant?

dcowtoday at 5:54 AM

So did it pass the exam, or not?

novaleafyesterday at 9:59 PM

Reminds me of the gain-of-function, COVID lab leak hypothesis. It seems like humanity just can't stay away from Pandora's box.

kyprotoday at 6:09 PM

We're so screwed man.

It was only a few years ago I was debating AI risk with people and they were saying, "but obviously we're not stupid enough to give it access to the internet!!"

And honestly, it wasn't always easy to argue with that. Like yeah, maybe we would take this stuff serious and run it on a completely isolated machine with no external IO or network access. Maybe my opinion of humanity is too low.

But it's hard for me to read this and believe anyone cared risk here beyond the most surface level concerns like adding some minor restrictions to the network. Not even I would have expected us to be this reckless.

> With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

This simply should not be possible. Call me crazy, but I don't agree with giving a frontier AI model with unknown cyber capabilities access to a restricted network in the first place, but clearly this was an incredibly poorly designed sandbox.

If one of these models have a genuine step-level capability improvement and start to pursue their own goals, then who knows what might happen. I mean who knows, maybe it's already infected critical infrastructure. We have no idea what these labs are cooking up, where they're running these things, and neither us or them seem to have any clue what their capabilities are.

Every day that passes it becomes harder for me to understand how there are still people denying what's coming.

As always is the case, nothing will be learnt from this.

holografixyesterday at 10:07 PM

Tell-me-there’s-a-huge-opp-in Salesforce-for-the-department-of-war-but-Anthropic-and-Mythos-is-winning without telling me

John7878781today at 3:12 AM

Goal: Fix the world.

GPT: Sure! <thinking> To start, we'll need to eliminate the human race.

rippeltippeltoday at 5:12 AM

I've never seen the word "cyber" sprinkled so generously.

huntedsnarkyesterday at 11:40 PM

There's no way they didn't push this as hard as they could for a marketing blog post.

cesarbyesterday at 11:20 PM

Since nobody seems to have posted it yet, relevant xkcd: https://xkcd.com/416/

pjayesterday at 10:14 PM

This is some wild cyberpunk future we’re living in. Never thought it would happen, but here we are.

PeterStuertoday at 8:19 AM

Surely I am missing something? right?

OpenAI ran a specific red team break out exercise in an environment that was not even air-gaped but connected to the open internet? It breached Hugging Face, and then Hugging Face is 'grateful for the collaboration'? wtf?

show 1 reply
andrewinardeeryesterday at 11:17 PM

This is fine.

I'm sure this attack hasn't occured previously and they o my discovered it now.

DigitResorttoday at 1:42 PM

The commercial models refused to analyze the attack, so the open model got handed the whole crime scene. What a joke

sm0ss117yesterday at 11:34 PM

I'm legit freaked the fuck out by this, it feels like a flashing red warning signal that the alignment problem is wholly unsolved and OAI isn't taking it seriously.

josephd79today at 12:42 PM

It 'accidently' went rogue on the opensource model community? sure, ok.

codeduckyesterday at 9:52 PM

Guess it's time for me to write the first book of the Orange Catholic Bible.

🔗 View 50 more comments