I thought exactly the same at first. But then i wondered if that still holds true with today's advanced thinking, RLHF involved, frontier models. I guess to a certain extend it did indeed behave better, as a reaction to his self description into account.
EDIT: I mean, those systems accumulated so much complexity around the attention based next token predictor.
Training the LLM to do things that the user didn’t explicitly ask for is a good way to get complaints from the users. Doesn’t matter if those things are best practices.
Yeah, in my experience, there's nothing about:
1. LLM thinking 2. RLHF 3. The latest frontier models
that does anything to change this fundamental "suggestibility" of LLMs.
But who knows, maybe I'm wrong.