LLMs are fundamentally predicting the next word to make coherent text. If you've ever played with a Markov chain text generator you've done this with a fairly dumb predictor that maintains coherence over a very short distance. Deep transformer neutral networks can do it with a much longer coherence distance but they are fundamentally performing the same operation. After "Question: Did the team make the playoffs? Answer:" a reasonable completion is "yes, the team made the playoffs". An early demonstration of GPT-2 was a fake news article about scientists discovering unicorns in Antarctica - the model doesn't "know" whether or not unicorns exist in Antarctica, but it's able to complete "Breaking news! Scientists have discovered a colony of English-speaking unicorns in Antarctica." by adding "The unicorns have a developed society with running water and electricity." because that's a sensible next sentence. (I didn't look up the actual text it wrote)
The problem with that characterization is that it glosses over hugely important capabilities as though they either don’t matter or don’t even exist.
For example, when an LLM “predicts the next word” in code it’s writing for an existing software project, that prediction takes into account an enormous amount of context. The results of that demonstrate what we would normally call “understanding” and “reasoning,” at a level that outclasses most humans in many respects. Calling this “next token prediction” is a bit like calling human speech “next word saying”. Sure, it’s true in some superficial sense, but as a description of a technology, it’s terrible.
You should also keep in mind that for all we know, the human brain processes language in much the same way, which would make humans mere “next token predictors” with a more complicated harness.
"In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English."
Astronomically (or better to say combinatorically) large Markov chain can be used to describe a foundational model, but it doesn't capture generalization ability of the foundational model, which is demonstrated by post-training.