logoalt Hacker News

ben_w • yesterday at 5:22 PM • 0 replies • view on HN

If they were totally unaligned, the GPT series would have never gotten past being autocomplete.

Literally all instruction following requires at a minimum alignment with attempting to implement those instructions.

We can argue about e.g. morality or law obedience on top of that*, but the general point is absolutely not avoidable.

* my position is that this tool is far too likely to metaphorically explode in the user's hands for companies to wash responsibility off on users: if OpenAI had released the model which did the HuggingFace attack, at a minimum thousands of random people (not all of whom would even be developers) would have issued instructions each with similar consequences.