I don't see why this is different to a careless developer allowing an agent to run rm -rf. I recognize the different angle with the exploits but boy wasn't this the exercise with ExploitGym?
Similar to how the basic thought "nobody gives you something for free" protects you from being ripped off in many situations we should apply "no AI company tells you about precious internals for transparency". It's stupid marketing and it's baffling to me how people give them any credibility.
Because “rm -rf” is a known, explicitly provided-in-docs-and-training command.
It is fundamentally different capability than “identified and chained multiple previously unknown exploits in order to bypass restrictions”. It’s even worse when/if the primary objective of this activity was to cheat on what it was doing.
It’s a foundational alignment issue, not a task-level result-alignment issue. Ie, “cheating” is fundamentally bad (when you have what are effectively rules of engagement), whereas deleting a directly is a thing that is correctly done sometimes (even if this invocation was a mistake/incorrect)