> We're copying morality and instruction following that seems to work on humans without really understanding why it seems to work on humans
To me it makes more sense to leave the models "unaligned" and leave it up to the operator to manage the morality of what they ask it to do. Besides, only humans can be charged with a crime.
If they were totally unaligned, the GPT series would have never gotten past being autocomplete.
Literally all instruction following requires at a minimum alignment with attempting to implement those instructions.
We can argue about e.g. morality or law obedience on top of that*, but the general point is absolutely not avoidable.
* my position is that this tool is far too likely to metaphorically explode in the user's hands for companies to wash responsibility off on users: if OpenAI had released the model which did the HuggingFace attack, at a minimum thousands of random people (not all of whom would even be developers) would have issued instructions each with similar consequences.