logoalt Hacker News

gwerbinyesterday at 10:59 PM3 repliesview on HN

But it literally is making a prediction based on the previous X tokens, it's just that X is huge and there is a proportionally huge number of parameters in the token generation function.


Replies

amlutoyesterday at 11:40 PM

It has nothing to do with X being huge. In fact X might be quite small.

danielmarkbruceyesterday at 11:28 PM

If you are going to say "literally", then what is your literal definition for the word "prediction" ?

wat10000yesterday at 11:15 PM

The discourse around this is annoying. A big group of people use "next-token predictor" to imply that LLMs aren't capable of anything interesting. Another big group of people opposes the use of "next-token predictor" because of that implication. But that fight isn't about the "predictor" language at all.

The linked article makes a good point: a substantial chunk of the training does not consist of "here's a bunch of tokens, here's the next token, learn that." But all the comments want to turn it into a referendum on the goodness of AI.

show 4 replies