logoalt Hacker News

cjonasyesterday at 6:55 PM1 replyview on HN

All it would have taken is someone to peak at the output tokens during the run and it would have been obviously the test had gone off rails.


Replies

ACCount37yesterday at 7:15 PM

Ha. As if.

1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent.

2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this.

3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" from "concerning AI weirdness" is not a trivial task.

show 2 replies