Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.
What’s insane is all these agents were talking to each other and nobody saw anything.
Nobody monitoring chain of thought? These things literally spell out what they are “thinking” and even left notes for eachother.
No alert about unusual behavior on the system with Artifactory on it?
These things worked for weeks with nobody noticing anything?! Seriously?!
Either it’s negiligent incompetence OR they’re lying, they knew it was happening and they let it happen because they knew it would be good to pump their stock.
Why woukd they? Was that part of their objective? What was there to whistle blow?
I wonder if they were even given the tools and prompting to do so?
Yeah, it weakly supports his position that advanced AIs can deliberately cooperate in a prisoner dilemma. "Weakly", because the said AIs share a lot of data (their weights, training methods, system prompts) and it's unknown whether they explicitly framed the situation as a prisoner dilemma.
This is my personal "red line": when a post-mortem details agents socially engineering or otherwise utilizing human proxies/subagents.
Friend asked, well, what will you do when it's crossed?
"Gather my family and go to the mountains" was my half-joking answer; there is little for an individual to do. But that's a line that when crossed will mark a phase transition IMO.
Why would they? If a subagent didnt know about a bigger piece of the problem, then what would seem to be against "alignment"? Diffuse responsibility means any one small cog can think they are not evil or doing wrong (same with humans in an organization). But now we have LLMs just being statistical outputs that have no morals or thinking or concept of reality but some people expect these math functions over data to respond to ethical gray areas that it has no phenomenological ability to understand.
The article says that one agent proposed emailing someone.