The same thing for both: Goodhart's Law.
That doesn’t apply in this scenario. If CEOs start trying to emulate ethics in order to avoid being 86’d, we still get a good outcome.
An issue arises when the metric isn’t a good proxy for the property being measured. But ethics would not be a simple numeric target. The comment above talked about a “score”, but the question is what goes into that score. It would need to be a list of items related to topics like discrimination, advocacy of inequality, tendency to circumvent regulations, etc. Set it up correctly, and even a CEO who’s willing to “fake it” would end up being better than most major company CEOs in e.g. finance, tech, or healthcare today.
EAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.”
That gives you a three-layer picture:
Task objective: Did it accomplish what we asked?
Acceptability constraint: Did it avoid unacceptable ways of accomplishing it?
Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint?
I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.