logoalt Hacker News

nradovyesterday at 9:04 PM1 replyview on HN

It's not technically possible to prove model alignment. At best you can establish a probability of alignment within certain constraints.


Replies

madroxyesterday at 9:48 PM

"Prove" is a shorthand because of course it's a stochastic process. My point is that I would expect ethics to be part of the training and eval criteria.