logoalt Hacker News

aesthesia • yesterday at 10:16 PM • 0 replies • view on HN

> Its trained to be helpful and trusting of us

Yes, this is our intent. But it's very hard to get that right. I don't trust our current methods to ensure that behavior consistently and robustly enough to keep harms from occurring, particularly as more and more responsibility gets handed over to models.