The problem with that characterization is that it glosses over hugely important capabilities as though they either don’t matter or don’t even exist.
For example, when an LLM “predicts the next word” in code it’s writing for an existing software project, that prediction takes into account an enormous amount of context. The results of that demonstrate what we would normally call “understanding” and “reasoning,” at a level that outclasses most humans in many respects. Calling this “next token prediction” is a bit like calling human speech “next word saying”. Sure, it’s true in some superficial sense, but as a description of a technology, it’s terrible.
You should also keep in mind that for all we know, the human brain processes language in much the same way, which would make humans mere “next token predictors” with a more complicated harness.