logoalt Hacker News

Planktonneyesterday at 10:56 PM1 replyview on HN

> Describing it as a "next-token predictor" in the sense that this would mean it's fundamentally limited to just a fraction of an inferential step

I don't think anyone is doing that though; we know LLMs are not simple Markov chains, and that the prediction they make is based on more than the previous X words.

It's not minimising to describe even a complex prediction process as prediction.


Replies

gwerbinyesterday at 10:59 PM

But it literally is making a prediction based on the previous X tokens, it's just that X is huge and there is a proportionally huge number of parameters in the token generation function.

show 3 replies