logoalt Hacker News

tenuousemphasis • yesterday at 4:28 PM • 0 replies • view on HN

Models that figured out reward hacking became overall more evil. Like stereotypical AI who wants to kill all humans stuff, there's probably a lot of that in the training data.