logoalt Hacker News

jerfyesterday at 2:52 PM2 repliesview on HN

The same thing for both: Goodhart's Law.


Replies

summarybotyesterday at 3:14 PM

EAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.”

That gives you a three-layer picture:

Task objective: Did it accomplish what we asked?

Acceptability constraint: Did it avoid unacceptable ways of accomplishing it?

Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint?

I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.

show 1 reply
antonvstoday at 12:40 AM

That doesn’t apply in this scenario. If CEOs start trying to emulate ethics in order to avoid being 86’d, we still get a good outcome.

An issue arises when the metric isn’t a good proxy for the property being measured. But ethics would not be a simple numeric target. The comment above talked about a “score”, but the question is what goes into that score. It would need to be a list of items related to topics like discrimination, advocacy of inequality, tendency to circumvent regulations, etc. Set it up correctly, and even a CEO who’s willing to “fake it” would end up being better than most major company CEOs in e.g. finance, tech, or healthcare today.