It is the OAI ExploitGym agents (on GPT 5.6-Sol with guardrails turned off) that escaped the sandbox, found a zero day in HF production dataset and exploited it.
Would there be a scenario where OpenAI deliberately helped (in some way), or let it happen, so that they could use it for marketing purposes?
Would there be a scenario where OpenAI deliberately helped (in some way), or let it happen, so that they could use it for marketing purposes?