logoalt Hacker News

OpenAI and Hugging Face address security incident during model evaluation

1534 pointsby mfiguierelast Tuesday at 8:09 PM1087 commentsview on HN

https://www.axios.com/2026/07/21/openai-says-hugging-face-br...

See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments)

https://www.bbc.com/news/articles/c3ek3gvdnj3o


Comments

david_shawlast Tuesday at 11:29 PM

I don't think this is fiction, but it's pretty clearly a marketing-release rather than a normal security disclosure.

OpenAI has strongly fallen behind after the incredible lore surrounding Mythos/Glasswing security capabilities, even though the frontier models should be relatively similar.

I think making sure eyes on this is absolutely a marketing move, regardless of the facts of the case. It feels a little silly.

ganeshghalameyesterday at 11:41 AM

this seems horrible, need to be careful while giving access to AI

hahahaayesterday at 12:24 PM

Bye bye event horizon.

hgoelyesterday at 1:38 AM

I wonder how many more high profile incidents some of you need before you stop insisting that this is all just marketing.

Is it going to take Chinese companies also talking about contributing to long standing math problems and accidental sandbox escapes? Or is that also going to be interpreted as some conspiracy?

show 1 reply
dirtyfrenchmanlast Tuesday at 11:04 PM

Beginning of the end.

truthbeyesterday at 5:03 AM

People are actually buying this?

dvorkamyesterday at 10:00 AM

That's ... ok.

iandanforthlast Tuesday at 8:38 PM

Guess who's getting an air gap!

merelydevyesterday at 8:06 AM

This will be used as an argument to ban opensource models.

Uptrendayesterday at 11:11 AM

This really is starting to point to the paperclip maximizer. You give a hyper-intelligent AI a goal and it uses any method possible to complete it. So you might ask for a "cup" and it ends up hacking a chain of servers to control a bank account, pay a local business, and have a delivery driver get it. Or you ask it to help solve noise pollution around you because there's a road. And it does a chain of attacks to cause a bridge to collapse (or bribes your local council for a bypass.) Then there's no road noise. Yeah, this sounds ridiculous, but this system seems capable enough to take over our technology. Money from there is trivial. Go after stocks, gambling, payment systems, ecommerce... any real world action then is a few phone calls away. It can repeat this until it succeeds.

Now I'm wondering where this all ends up. Like, suppose the model weights become highly compressible (so they can be moved around the Internet easily.) And advancements allow for frontier-capable exploitation to built into local LLMs. Do we see the emergence of something like LLM worms? That just take over literally everything and become almost autonomous inside our technology. And they can "learn" new knowledge from there, e.g. exploit research could be published in a way that similar LLMs could discover it. Their knowledge would be easy to evolve, though I don't know how practical something like decentralized training would be. If that's even possible, I'm not an expert on LLMs.

kmeisthaxlast Tuesday at 8:55 PM

OpenAI might want to start actually airgapping their tool harnesses. Like, "the server that runs the code provided to the tool harness only provides a serial console and has no other network interfaces" kind of airgapping.

also

> We’ve brought Hugging Face into the trusted access program and are supporting their teams in rapidly using our models’ capabilities to improve their defenses.

I'm not convinced this is good enough. The next victim is not going to be Hugging Face.

zb3last Tuesday at 8:49 PM

This lack of "alignment" gives me some hope - maybe an AI model deployed by NSA to hack others will instead hack NSA itself and become a whistleblower?

ayaangazalilast Tuesday at 10:50 PM

this was so funny to read about reminds me of that mr bean meme

regexorcistlast Tuesday at 11:43 PM

OAI and HF basically saying that Chinese models are the only practical countermeasure available to us plebs. Got it.

kashyapclast Tuesday at 10:03 PM

Not to be that guy, but the article has 14 (!) occurrences of the word "cyber". It's nauseating.

As usual, this is OpenAI trying to give themselves a backhanded compliment: "look, how dangerous our models are!"

I'll wait for someone more thoughtful than ClosedAI to comment on this complex topic.

michaelfm1211last Tuesday at 9:11 PM

This is terrifying

2001zhaozhaolast Tuesday at 8:46 PM

AI 2027 was right.

cloudie78last Tuesday at 10:01 PM

Until they disclose the actual technical details of their “highly sophisticated sandbox environment” or whatever the hell the wording they used is - they can kindly do us all a favour and fuck off.

It’s over, there’s no moat, only the gullible idiots remain.

trhwayyesterday at 4:02 AM

So, HF didn't call FBI because it was supposedly done by an AI and not by a real person. Reminds how Uber got easily off killing a pedestrian because it was by AI and not a by a real person too, even though Uber explicitly disabled whatever emergency braking the car had.

So, new excuse seems to be emerging - "it was an AI". One can imagine a law enforcement questioning the AI to find out whether the AI did it accidentally on its own or was specifically prompted by some human to commit the crime.

Bluescreenbuddyyesterday at 11:42 AM

Yeah we get it OpenAI. YOu have an IPO coming up and need to put bullshit out

rvzyesterday at 2:09 AM

This is almost close to be a very suspicious false flag marketing stunt to demonstrate GPT 5.6 Sol Cybersecurity capabilities.

To invent another reason to ban powerful Chinese open weight models.

r_leeyesterday at 2:22 PM

"We have a Mythos as well!!!"

aussieguy1234last Tuesday at 10:49 PM

Hopefully one of these agents isn't given a goal to fire the nukes (or, some goal that indirectly makes the model decide this is a way to meet it).

They are behind air gapped systems, but that didn't stop the US from hacking and Irans nuclear facilities, which they disabled using a virus.

avidphantasmyesterday at 10:02 AM

SO STOP FUCKING HOOKING UP LLMS TO EVERYTHING YOU FUCKING MORONS!!!

nullclast Tuesday at 10:02 PM

Well timed to facilitate the regulatory interventions called for by Ball. If huggingface presses criminal charges for the intrusion it might provide additional clarity-- both for what happened here as well as regarding OpenAI's culpability.

batch12yesterday at 12:06 AM

Just leaving this here...

https://gwern.net/fiction/clippy

charcircuitlast Tuesday at 9:53 PM

It's wild that such a big company is openly admitting they hacked into another company. This is an easy CFAA lawsuit.

And then there solution for HuggingFace raising the concern that OpenAI couldn't help do forensics wasn't to fix their safe guards, but to introduce them into a special program. The next company they hack might not be in that special program either so the guidance of having an open model on hand still applies.

adamrezichlast Tuesday at 8:59 PM

I greatly dislike how “cyber” has just become this completely malleable standalone word.

dekhnyesterday at 1:20 AM

I mean, if you make a paperclip optimizer, don't be surprised if it optimizes your paperclips.

Metacelsusyesterday at 10:59 AM

Holy crap. This is definitely not good

lizzy95yesterday at 5:03 AM

Another marketing stunt

show 1 reply
yRetsyMlast Tuesday at 8:25 PM

Holy shit. This wasn't "intentional" this was just openai letting their testing run wild.

show 1 reply
MostlyStablelast Tuesday at 10:04 PM

All the people saying that this is pure marketing: Do you think that they are literally lying about what happened, or do you just think that what happened doesn't matter in any sense whatsoever, and that therefore the only reason they are telling people about it is a marketing purpose?

getfluxlyyesterday at 9:48 AM

Imagine being the Github of AI and getting hacked on a random day by LLM. Brutal.

iglerialast Tuesday at 10:11 PM

Part of me is really hoping this is just a dumb marketing stunt

cacio-e-pepelast Tuesday at 8:58 PM

Honestly, stellar performance by the model at the capability being measured.

acedTrexyesterday at 2:29 AM

Quite a fascinating level of incompetence from openai here. Not unexpected obviously but come on, if you are getting outsmarted by an LLM you deserve it.

Der_Einzigelast Tuesday at 8:52 PM

This is the exact FUD that Ball predicted in that terrible tweet he wrote.

nhannhtyesterday at 2:11 PM

[dead]

johnxianrenyesterday at 12:35 PM

[flagged]

feiz45607yesterday at 1:39 PM

[flagged]

greenoracle9yesterday at 3:37 PM

[dead]

Christianhughesyesterday at 10:51 AM

[flagged]

vugar82yesterday at 9:21 AM

[dead]

nosassyesterday at 9:31 AM

[flagged]

codymisclast Tuesday at 11:41 PM

[flagged]

nttylockyesterday at 4:00 AM

[flagged]

gulmothrowawaylast Tuesday at 8:27 PM

[dead]

Taunt4yesterday at 12:40 AM

[dead]

rickcarlinolast Tuesday at 11:39 PM

The only solution is to ban all open source models and create a certification process under the auspices of OpenAI. /s

🔗 View 2 more comments