Yesterday I came up with an idea that I sent to some researchers at the different AI labs via email: Rather than train the model on one score, track two scores. The first score is the Short-term-objective-score (STOS) and the other, more important one, is the EAOS Ethically-aligned-outcome-score. Every trajectory can be evaluated on whether or not it has a high enough EAOS to be considered acceptable. If the model does some task and has a very high STOS but very low EAOS, like modifying game code to win at a game rather than playing by the rules, it is unacceptable. Models going forward must all have an ethics evaluation in tandem with objectives evaluation, and only when the ethics value is high enough should actions be considered successes.
There's no way to make an EAOS score automatically. If we had that, that's the whole fix. Just reject answers with low ethics numbers.
Whats the definition of EAOS though who's ethics? Greek-Roman, Western, Islamic, Buddhist, Hinduism, Human rights (western values)..
1) that calculation is being done even without a cost function 2) trying to make a cost function for EAOS is impossible 3) gaming/goodharts. There's no good solution. Best we can do is push for decentralization, open source, regulatory capture, etc. Of course third party metrics might be good, especially if there's tons of them with well-documented rationale.