I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.
It should know which actions are ok and which aren't. Maximizing paperclip production should be within your factory (or talk to the boss about opening more), not world domination or nuclear war. Solving problems shouldn't involve hacking other systems or escaping a sandbox.
> I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.
> It should know which actions are ok and which aren't.
It's worse than that:
They do know, we can see them write down notes that certain actions are forbidden.
They then go off and performs the actions anyway.
My expectation for the cause? Helpful vs harmless: you can pick anywhere from one to the other, but you can't get both at the same time. The models are trained to do what the user tells them to do.
Just look at all the pushback the model makers get when they put in guardrails:
- user matheusmoreira, here, 13 days ago: https://news.ycombinator.com/item?id=49678048This user will not be alone; their preferences, and similar from others like them, will form part of any RLHF-style training.