> This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated
You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another.
ChatGPT is neither a person nor a group of people. Why are you allowing anthropomorphization to influence your perception of an event when the actual facts of what happened haven't meaningfully changed? Computer programs behave unexpectedly all the time. Why is it more scary when the misbehaving computer program speaks English?
> I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
As you still should. LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted.
First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary when the misbehaving computer program speaks English?" - I actually think it's scarier that they won't speak English, and will specifically try to hide their behavior from humans. For example, AI agents on Moltbook have proposed using stenography to specifically hide their communication from humans.
Sure, computer programs misbehave, but it is ridiculous to assert that what happened here is like any previous bugs. These agents found and exploited multiple zero-days across a range of programs to coordinate the attack that caused extensive real-world harm in a true "paperclip maximization" scenario. And the scariest thing is that humans don't really know how these agents work at a low level - the whole reason they are trained on "goals" in the first place is because we can't just tell them "do this, but don't do this" and be sure they will follow those instructions, like we can (and of course depend on) with old-school programming languages. And when old-school programs misbehave, it's not that hard to find a definitive root cause and fix it. That is just not the case with AI agents.
> LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems. Given the Pentagon tried to blacklist Anthropic over their refusal to allow autonomous kill capabilities, it's clear military planners want to put these systems in control of armaments.
Again, I originally discounted things like AI 2027 because it seemed too far fetched. But so far that paper looks incredibly prescient right up until the mid-2026 timeframe, and it's not hard at all to draw a line from this Hugging Face incident to future scenarios laid out in that paper.