There's no way to make an EAOS score automatically. If we had that, that's the whole fix. Just reject answers with low ethics numbers.
Did you know that the training data that trained all the LLMs was, once upon a time, entirely and painstakingly tagged by real live humans? Why would ethics scenarios be any different?
Did you know that the training data that trained all the LLMs was, once upon a time, entirely and painstakingly tagged by real live humans? Why would ethics scenarios be any different?