It's not technically possible to prove model alignment. At best you can establish a probability of alignment within certain constraints.
"Prove" is a shorthand because of course it's a stochastic process. My point is that I would expect ethics to be part of the training and eval criteria.
"Prove" is a shorthand because of course it's a stochastic process. My point is that I would expect ethics to be part of the training and eval criteria.