Respectfully, go build one, including doing RLHF and RLVR. Those phases generate lots of tokens, then get scored on the entirety of the output, then optimize based on a scoring of that output. It doesn't check a "prediction" against what was actually "next" in data, because there isn't any "next token" data it's training on.
> It doesn't check a "prediction" against what was actually "next" in data
Literally no one here is claiming that it does. This is one of the many flaws in the article.