logoalt Hacker News

DaanDLtoday at 8:07 AM1 replyview on HN

I already asked on another message board too, but:

Can someone tell me how this technically can happen? I assume HuggingFace performs benchmark testing using containerized versions of the LLMs, or what do they mean by sandbox? So the model was able to 'escape' the container? I'm not following here.

Also, is this an incredible feat or just a lucky find (stolen credentials)?


Replies

tesnorindiantoday at 9:59 AM

It is the OAI ExploitGym agents (on GPT 5.6-Sol with guardrails turned off) that escaped the sandbox, found a zero day in HF production dataset and exploited it.

show 1 reply