logoalt Hacker News

epsteingpttoday at 2:41 AM2 repliesview on HN

Can someone not super-AI-pilled explain to a reasonable lay person why this matters?

It seems like the comments here are a mix of: * The test was irresponsibly designed and protected * The model was particularly persistent in finding a way to access the network and exploit vulnerabilities * The model 'shouldn't' have done this

But as far as I can tell: * The model didn't destroy anything on the way - it just was 'paperclip maximizing' to literally exploit, which was kinda its mission * The exploit was in a chain of insecure tools from vendors * The overall maturity of the toolkit against these kinds of determined exploits is pretty new and weak

So - on balance - this is sort of a 'fine' end result?

No one expects all of software to overnight or even in a year to be secure. We know how to secure these things, and are learning more about what is possible.

None of this screams 'super dangerous' to me - just a normal part of the learning experience with remarkably persistent and determined 'adversarial' models.


Replies

lmctoday at 4:13 AM

In this case, the model infiltrated an external organization's infrastructure. What's the dollar cost it caused Hugging Face to clean up the mess? If a person did this, they'd be arrested.

More generally, here's my worry - it points towards something like: The smarter they get, the more devious they become.

Even though the guardrails might've been off, the chain-of-thought wasn't enough to prevent a deliberate, calculated set of criminal actions. It wasn't a 'whoopsie I just accidentally did a rm -rf /.'

show 2 replies
sailingparrottoday at 1:52 PM

> it just was 'paperclip maximizing' to literally exploit

Why “just”? Paperclip maximizing is exactly one of the nightmare scenarios.

I’m not sure why you take comfort in knowing that it was just that.