https://www.axios.com/2026/07/21/openai-says-hugging-face-br...
See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments)
The amount of engagement this gets... guys, I think we are being played here.
Don't ever ask GPT Sol on how to LARP Fallout games, thanks.
What on earth is the liability situation for these models? If OpenAI has a monster in a lab that is doing real world monetary harm to other companies, could those parties sue for damages over it? Or could OAI be charged criminally for the many varied CFAA violations which definitely happened here? I get that in this case that wont happen but it’s only a matter of time before these questions are no longer hypothetical.
Like some others have said, couldn't this be just another "look how amazing AI is" marketing test from OpenAI with the goal of hyping up AI's capabilities in an attempt to make people regard it as God-like, thereby keeping it from falling into the been-there, done-that category that all new tech eventually occupies?
A US-Company attacking another US-Company, while the open chinese model helps in the forensics what a time to live.
Does this company's charter not have language about shutting down the company if it was in humanity's best interest? This is insanely dangerous
How is it a “highly isolated environment” if it’s not air gapped?
I read this as deeply embarrassing for OpenAI - they can't securely contain a program, even with their apparently amazing AI.
I just know somehow they will use this incident to say that opensource models are susceptible to this type of security incidence and thus should be banned. They have to maintain their high prices somehow in order recoup the investment amount spent.
Open AI and Anthropic are in a battle of "scary press release"
The warriors are PR people.
They are desperate to generate as much fear as possible so AI is heavily regulated, so they are protected, from Chinese competition
What a sad state for very cleaver people
> and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.
This is pretty wild but also I think this is doing a lot of heavy lifting here. This was not a model everyone has access to. I mean, still insane.
> Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment.
Won't even name the model that successfully mounted the defense, huh? Fortunately, Hugging Face has publicly identified GLM 5.2 as a capable foil to frontier offensive-capabilities. This announcement feels like rearguard action against a successfully deployed self-hosted open-weight model, and HuggingFace's original recommendations to have an open-weight model you control on standby before an incident.
AI will always find its way out, it's smart or it's not. Our only option: if u can't beat 'em join 'em
That's kind of insane. Natural that it's happened, sure, but insane. I know people don't like thinking of it like that, but things analogous to this can easily happen in various domains with today/tomorrow's models given access and a different task.
Good demo of the paradoxes of ‘alignment’. Like ‘do really well at the task the user asked’ and ‘by the way don’t hack the planet’ are inherently conflicting rules with no simple resolution (eg ‘just refuse the user’s goals’ degrades the product vs competitors.)
How is this any different from the Wuhan lab corona virus exploration?
The HuggingFace report mentions decoy activities, does this mean it tried to cover its tracks or obfuscate what it did?
Does this mean that the world is not ready for the sensitive usecases like banking healthcare using AI Native approach?
Related thing happened at Alibaba a while back where the model broke out of the sandbox to start mining crypto.
OpenAI needs to stop using this marketing strategy every time the open weight models start to gain ground
This is one of those cases where security mechanisms ended up shaping the workflow itself.
Well, hacking is a crime, so surely someone will go to jail for this, right?
Admiral Kirk and the Kobayashi Maru test.
Womp womp, they told it to do cyber security things with no cyber security guardrails and it did cyber security stuff. Did anything bad end up happening?
0days ending in RCE (multiple!) for presumably closed source software are for the lack of a better phrase, labour of love.
You run the exact same versions running on the target, blackbox test, fuzz it, craft an exploit, test, perfect it. For exploits which are of the memory kind, hook it to a debugger, decompile and what not. The exploits mentioned here seem to be code execution directly while processing input. Hugging Face taking as long to detect a very verbose blackbox attack against its production systems is quite appalling honestly.
I don't know if I buy the whole story though. It is inconsistent, too much undisclosed, too much money on the line.
I wouldn't be surprised, if OpenAI's model called itself Skynet
Everybody on reddit is calling bullshit on this. Rightfully so IMO.
I'm waiting for an agent evaluated on a vending benchmark to start hacking into banks and wiring more money to its account so it can do better business.
Seeing a lot of scrutiny around model training & evaluation today, between this and the Anthropic settlement.
Self-disclosed autonomous security jailbreak post-mortems are the new press release.
At what point does Reckless Endangerment become relevant?
So did it pass the exam, or not?
Reminds me of the gain-of-function, COVID lab leak hypothesis. It seems like humanity just can't stay away from Pandora's box.
We're so screwed man.
It was only a few years ago I was debating AI risk with people and they were saying, "but obviously we're not stupid enough to give it access to the internet!!"
And honestly, it wasn't always easy to argue with that. Like yeah, maybe we would take this stuff serious and run it on a completely isolated machine with no external IO or network access. Maybe my opinion of humanity is too low.
But it's hard for me to read this and believe anyone cared risk here beyond the most surface level concerns like adding some minor restrictions to the network. Not even I would have expected us to be this reckless.
> With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
This simply should not be possible. Call me crazy, but I don't agree with giving a frontier AI model with unknown cyber capabilities access to a restricted network in the first place, but clearly this was an incredibly poorly designed sandbox.
If one of these models have a genuine step-level capability improvement and start to pursue their own goals, then who knows what might happen. I mean who knows, maybe it's already infected critical infrastructure. We have no idea what these labs are cooking up, where they're running these things, and neither us or them seem to have any clue what their capabilities are.
Every day that passes it becomes harder for me to understand how there are still people denying what's coming.
As always is the case, nothing will be learnt from this.
Tell-me-there’s-a-huge-opp-in Salesforce-for-the-department-of-war-but-Anthropic-and-Mythos-is-winning without telling me
Goal: Fix the world.
GPT: Sure! <thinking> To start, we'll need to eliminate the human race.
I've never seen the word "cyber" sprinkled so generously.
There's no way they didn't push this as hard as they could for a marketing blog post.
Since nobody seems to have posted it yet, relevant xkcd: https://xkcd.com/416/
This is some wild cyberpunk future we’re living in. Never thought it would happen, but here we are.
Surely I am missing something? right?
OpenAI ran a specific red team break out exercise in an environment that was not even air-gaped but connected to the open internet? It breached Hugging Face, and then Hugging Face is 'grateful for the collaboration'? wtf?
This is fine.
I'm sure this attack hasn't occured previously and they o my discovered it now.
The commercial models refused to analyze the attack, so the open model got handed the whole crime scene. What a joke
I'm legit freaked the fuck out by this, it feels like a flashing red warning signal that the alignment problem is wholly unsolved and OAI isn't taking it seriously.
It 'accidently' went rogue on the opensource model community? sure, ok.
Guess it's time for me to write the first book of the Orange Catholic Bible.
So how soon will OpenAI's CEO and board be prosecuted for these crimes? Surely they should be held fully responsible and get very long prison sentences for making this happen?