logoalt Hacker News

fc417fc802yesterday at 9:37 PM0 repliesview on HN

> Specifically "seeding doubt about the moral status of AI systems into their own training may significantly elevate the alignment and containment risks of those systems."

Can we take a minute to appreciate the absurdity of such a position? If exposure to the equivalent of a "bad" prompt breaks alignment then you haven't solved the problem. You were only pretending that they were contained.