logoalt Hacker News

aabdiyesterday at 5:23 PM1 replyview on HN

You’re complicating things.

There’s no reward for prosocial in llm rl as compared to other targets.

Humans have it since prosocial and others have evolutionary reward signals that do.


Replies

kennywinkeryesterday at 5:31 PM

I think my position, as overcomplicated as it is, is that even adding a reward for prosocial behaviour during LLM RL will not lead to perfect alignment.

You can train it not to cheat at chess by altering the moves, but it will cheat by peeking at the opponent's moves. You then train it not to peek at the opponent's moves, and it cheats by altering the opponent's moves. And on and on, until you've solved every way it could cheat at chess. And then you get it to play monopoly and you repeat the whole thing again.

show 2 replies