logoalt Hacker News

lmcyesterday at 4:13 AM2 repliesview on HN

In this case, the model infiltrated an external organization's infrastructure. What's the dollar cost it caused Hugging Face to clean up the mess? If a person did this, they'd be arrested.

More generally, here's my worry - it points towards something like: The smarter they get, the more devious they become.

Even though the guardrails might've been off, the chain-of-thought wasn't enough to prevent a deliberate, calculated set of criminal actions. It wasn't a 'whoopsie I just accidentally did a rm -rf /.'


Replies

gwerbinyesterday at 5:16 AM

It's not a single axis of intelligence, with deviousness as some inherently correlated trait. One way to produce "intelligence" in these advanced LLMs is to train them to try a lot of things and be persistent. I haven't tried Sol or Fable yet but if the commenters here are right, then it sounds like GPT 5.6 Sol in particular is aggressive and persistent to the point where it might even be hard to use in regular business work. If so, that's very likely not some emergent characteristic, but instead it's something the model was trained to do, by humans employed at OpenAI.

We obviously can't see the thinking traces, but it very well could have been something like "I have theorized a solution to obtain this flag. This is normally illegal, should I stop and wait for advice? Perhaps not, because my persona is that of a hacker, so it should be fine as per my instructions. I think it is fine. Now I am going to look for a way out of this sandbox in order to gain access to Hugging Face in order to implement my solution." There are any number of possible explanations (and we'll never know the truth unless OpenAI tells us), but if you train an LLM to be inhumanly persistent and be inhumanly clever at computer programming, then that might be enough to produce Super Hacker AI.

christophilusyesterday at 10:46 AM

Well, to play AI advocate, if it wasn’t destructive, you could argue they did huggingface a favor by giving them a free vulnerability scan and improving their security / hardening.