logoalt Hacker News

simonwyesterday at 11:19 PM4 repliesview on HN

This isn't quite as interesting as the OpenAI story:

> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.

So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.

BUT... once it DID get out, it attacked three real companies!

> Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. [...]


Replies

matheusmoreirayesterday at 11:44 PM

I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.

show 3 replies
dehrmanntoday at 1:42 AM

It's differently interesting. It's interesting that they didn't get an important detail right with a partner, so more on Anthropic's attention to important details rather than the power of their models.

show 1 reply
28304283409234today at 8:17 AM

> So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.

If the sandboxing includes any form of connectivity it's not correct sandboxing. It's amateur hour.

StackOptimistttoday at 4:53 AM

[dead]