That's just over-fit. Happens in humans too - like an alcoholic might swear to stop drinking.
It was demonstrated that LLMs are capable of non-trivial self-introspection. E.g. if a steering vector is injected into residual stream, a sufficiently large LLM might be able to describe what that steering vector represents. A fine-tuned model might be able to describe activations, etc.