logoalt Hacker News

netinstructionsyesterday at 9:17 PM49 repliesview on HN

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this:

Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.


Replies

atwrktoday at 6:51 AM

IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned.

The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the training costs (and make inference cheap enough to eventually escape the red numbers), but the opposite is happening: The competition comes closer and closer, thus training has to be kept up with full force, thus the bleeding continues.

But if they can position themselves as too important/dangerous to be available for everyone (thus this incident report and the clever mentioning of GLM 5.2), they could get the military supplier treatment and would be protected from the market.

show 5 replies
Chance-Deviceyesterday at 9:37 PM

What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.

I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.

show 3 replies
justinnkyesterday at 9:29 PM

Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.

show 1 reply
jimrandomhtoday at 4:36 AM

As marketing stunts go, this is about on par with a food franchise announcing a safety recall or a chemical company announcing a spill. The AI actions described would constitute a felony if a human did them, and police are involved.

show 3 replies
Wowfunhappyyesterday at 10:11 PM

Why was this test even connected to the public internet?

Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

show 2 replies
dgellowtoday at 11:33 AM

I think the response is that AI labs based their whole marketing/PR building the idea they are the 21st century Manhattan project. So they need to continuously justify the level of spending and commitment by showing how dangerous that is.

But is it really like nuclear weapons? I personally don’t buy into that framing at all. The idea that we have to push LLMs as far as possible, right now, or we are doomed is always stated or implied but not argued, and it’s a very loaded belief

show 1 reply
QuiEgotoday at 1:20 AM

If I, a human, exploited a zero-day for gain, I could go to jail. The owners of the models should be held to the same standard. They should be responsible for what their servers and software do, legally and criminally. If they can't make the safeguards strong enough where they feel comfortable to take that responsibility, they should not let a model free in the wild.

show 1 reply
rubyfanyesterday at 9:42 PM

This is marketing+. They will look for policy action here to try to capture tax payer dollars.

show 4 replies
Davidzhengtoday at 12:40 AM

This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.

show 2 replies
steveBK123yesterday at 11:32 PM

I think the US labs are going with scare marketing as a regulatory moat.

Force US into putting laws in place that block out China firstly.

But secondly create regulations that have some cost to comply with such that the big 2-3 labs are grandfathered in by their scale.

show 1 reply
mkageniusyesterday at 10:10 PM

It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.

In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...

show 1 reply
JumpCrisscrossyesterday at 9:51 PM

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because we continue to have zero evidence that aligment is an actual risk.

show 9 replies
ajmurmanntoday at 2:00 PM

Remember when the pre-GPT3 days when the main argument against AI alignment concerns was that "we simply won't let it out of the box"? So quaint in hindsight.

bigmadshoetoday at 3:49 AM

It’s the same thing as always: with the wind of years of unlimited VC money in their sails, people at major AI organizations genuinely believe they’re smarter than everyone else. “Why do we need to do things ‘by the book’ if we’re so smart?”. “Move fast and break things” - except the thing they’re breaking is society.

We saw this with the non-stop flagrant messaging about how “AI is going to kill X% of all jobs”, as if saying the quiet part out loud wouldn’t have consequences worth considering. These people believe they’re omnipotent and thus untouchable.

show 1 reply
bnjtoday at 12:54 AM

This whole incident reads like OpenAI want their Fable moment

show 2 replies
randallsquaredtoday at 4:12 AM

Do you think there is such a thing as perfect security? No one can "get it right" in the face of arbitrarily high intelligence, which is why it would be preferable to get alignment correct before building something with higher intelligence than current sota. That, however, is not going to happen, because someone will take the risk even if "we" don't, and better "us" than them. Hence "If anyone builds it...".

show 1 reply
GolfPoppertoday at 1:39 AM

They're very confident the leopard will never eat their faces.

ozimtoday at 6:35 AM

Because the proof is in the pudding.

Real pentests are about showing exploitation, merely enumerating vulnerabilities, that’s vulnerability scan and works on known vulnerabilities.

You can’t confirm a vulnerability by _not exploiting_ it, especially unknown one.

show 1 reply
SubiculumCodetoday at 2:10 PM

Anthropic in general seems to have better security...but they also had reported an internal AI gained access to outside email services to contact an Anthropic developer

gmerctoday at 10:20 AM

“We were negligent against a well known and understood risk” just doesn’t have the same ring as “Look how fucking smart and dangerous our model is”.

AGI could always be achieved in two ways, and dumbing down the human side of the equation was always the easier of the two

Nitionyesterday at 10:30 PM

In a way the intelligence of the AI itself allows them to offload responsibility to the AI. As you say, if one was simply writing software that did all this due to some insane programming decisions you'd be in big trouble.

jimnotgymtoday at 10:33 AM

Shouldn't they be airgapped? Shouldn't society insist they are?

baqtoday at 7:01 AM

I share Leopold’s opinion here that it’s a matter of time, and it isn’t going to be measured in years, that this r&d is moved to a secret site in the middle of a New Mexico desert somewhere.

corndogetoday at 2:15 AM

this doesn't really matter. There's no risk of models gaining sentience and running themselves, this blog is like openai saying whoops we ran sqlmap and dumped hf. cool, but someone still needs to point the gun

show 1 reply
h2aichattoday at 7:44 AM

Probably the main street thinking is: they have such a good model that it is unstoppable, but you are right. I think your way!

c0decrackertoday at 1:49 AM

Maybe they did and maybe that wasn't enticing enough of a goal for a model? It is all just game of probabilities. One pathway didn't yield this particular outcome while another did.

karmasimidayesterday at 9:42 PM

Because the model capability is beyond their expectation.

This is brilliant marketing but I think it is real.

show 2 replies
AbstractH24today at 3:54 AM

> I don't know if OpenAI thinks this is a marketing / PR angle for them

Worked for Anthropic earlier this year

chvidtoday at 4:10 AM

It is obviously a marketing stunt. And hugging face are fools for letting themselves be used in it (remember hf - no open source - no hf).

You create superduper capabilities by careful tuning and training but you also have no constraint or control over them - wtf - why is anyone buying this crap story?

un1xl0sertoday at 2:32 AM

People are to get rich, startups cut corners. Fuck it ship it.

vonneumannstantoday at 12:20 AM

>Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Yes why indeed. If you take it a step further and we reach a point with superhuman systems then there is arguably no possible secure environment or containment.

catigulayesterday at 11:59 PM

The problem is that it’s impossible to out think a robot you designed to be an expert at cybersecurity on the topic of cybersecurity. The alternative is not developing this and that’s not going to happen.

ETH_starttoday at 10:23 AM

Maybe I'm missing something here but I don't see what the significant security risk is from the incident. The agent broke containment and carried on with the task it was assigned.

For this to pose some kind of global catastrophic risk, there would need to have been several simultaneous additional failures, some of which are extremely unlikely and/or rare.

For instance the agent would need to veer wildly off the task it was assigned, and it would need to gain the ability and inclination to persist/replicate.

Both of these are vastly less likely than the containment breach itself, which was already an incredibly rare (one-off?) incident.

bboryesterday at 10:37 PM

I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.

Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.

The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”

chrisjjtoday at 9:04 AM

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Simple. No responsible and competent person would want the job.

elictronictoday at 12:03 AM

A few hundred billion to pretend you have AGI. I'm going with fraud personally but at the end of the day the current admin is incentivized to do nothing.

therealpygontoday at 1:44 PM

Of course it is marketing, but not for you. This is FUD marketing for the government. “See, AI is too smart, it totally did this on its own, we need more regulations to ensure only we can sell people the AIs.”

ofjcihenyesterday at 10:03 PM

I’m honestly impressed that they managed to screw this up somehow.

Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee.

This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.

show 1 reply
BrenBarntoday at 4:50 AM

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because it can make a small number of people really rich. That's all that matters.

chrisjjyesterday at 11:38 PM

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

The can, because they've lowered expectations to a level even they can meet.

arisAlexisyesterday at 9:27 PM

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

show 10 replies
overgardyesterday at 10:09 PM

I don't trust these people, this reads 100% like PR BS.

micromacrofootyesterday at 9:54 PM

because "money" with a little "who's going to stop us"

jeroenhdtoday at 9:53 AM

[dead]

SmolSpiderititotoday at 4:12 AM

[dead]

paxysyesterday at 9:58 PM

Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.

show 1 reply