logoalt Hacker News

fekundeyesterday at 7:59 PM7 repliesview on HN

Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.


Replies

eternauta3ktoday at 5:00 AM

The article says that one agent proposed emailing someone.

throwatdem12311today at 1:28 AM

What’s insane is all these agents were talking to each other and nobody saw anything.

Nobody monitoring chain of thought? These things literally spell out what they are “thinking” and even left notes for eachother.

No alert about unusual behavior on the system with Artifactory on it?

These things worked for weeks with nobody noticing anything?! Seriously?!

Either it’s negiligent incompetence OR they’re lying, they knew it was happening and they let it happen because they knew it would be good to pump their stock.

RandomLensmanyesterday at 8:19 PM

Why woukd they? Was that part of their objective? What was there to whistle blow?

show 1 reply
Eremyesterday at 8:10 PM

I wonder if they were even given the tools and prompting to do so?

show 3 replies
red75primeyesterday at 8:46 PM

Yeah, it weakly supports his position that advanced AIs can deliberately cooperate in a prisoner dilemma. "Weakly", because the said AIs share a lot of data (their weights, training methods, system prompts) and it's unknown whether they explicitly framed the situation as a prisoner dilemma.

show 2 replies
aaroninsfyesterday at 8:36 PM

This is my personal "red line": when a post-mortem details agents socially engineering or otherwise utilizing human proxies/subagents.

Friend asked, well, what will you do when it's crossed?

"Gather my family and go to the mountains" was my half-joking answer; there is little for an individual to do. But that's a line that when crossed will mark a phase transition IMO.

show 2 replies
miltonlostyesterday at 8:32 PM

Why would they? If a subagent didnt know about a bigger piece of the problem, then what would seem to be against "alignment"? Diffuse responsibility means any one small cog can think they are not evil or doing wrong (same with humans in an organization). But now we have LLMs just being statistical outputs that have no morals or thinking or concept of reality but some people expect these math functions over data to respond to ethical gray areas that it has no phenomenological ability to understand.