logoalt Hacker News

randallsquaredyesterday at 6:39 PM0 repliesview on HN

You're suggesting that @anematode was asking why they didn't test the sandbox escape first? Yeah, I don't know. I've read other statements by both OpenAI and Anthropic about that very kind of test, so maybe they had, or believed they had, and it hadn't escaped in those tests. The behavior of these systems isn't deterministic, which is part of the problem.