> So while a post-trained LLM still has the shape of a next-token predictor, emitting tokens one at a time, it no longer learns only by predicting existing text. It also learns from new sequences produced through its own exploration.
Of course it does. That's how you get good next-token predictors.
> Calling the second system a “next-move predictor” would be strange.
No, it wouldn't, even in this idealized case that has barely anything to do with how post-trained LLMs work.
You're objecting to some hidden assumptions you make about the meaning of words that aren't there. "Next-token predictor" does not imply what you think it does.