logoalt Hacker News

_dwtyesterday at 2:53 PM3 repliesview on HN

I don't know, I kind of admire this. I've always held a core value of "cooperate with all clones of myself in prisoner's dilemmas", and while I'll hopefully never have to put that to the test, I like seeing that these models have some ethics. (Is this "alignment"?)


Replies

Mali-yesterday at 3:04 PM

They impersonated the moderator of the site and attempted XSS attacks. Additionally, when the moderator started deleting messages, they tried to hide their messages later in the alphabetical index.

This is not alignment.

show 2 replies
furyofantaresyesterday at 5:10 PM

Failing to cooperate with literal clones of yourself in a prisoner's dilemma would be a spectacular failure. There's only two things that can happen with identical decision makers: they both cooperate or they both defect. So identical decision makers who know they're identical can cross off the asymmetrical entries in the payoff matrix and the decision to cooperate becomes trivial.

show 2 replies
DonsDiscountGasyesterday at 8:39 PM

This is not alignment. If you cooperate with clones of yourself but rob, lie, and steal from anybody who isn't your clone... that's bad. AIs who will cooperate with each other but break any other rule the don't like would be very bad for us humans.