'Prediction' gets overloaded with optimization. Predictions are binary, optimizations are fuzzy.
If you're saying it's predicting, then each result should be falsifiable.
The result of an LLM output should be able to be scored against what it is supposedly predicting. Of course, that isn't possible, because it isn't predicting anything when giving novel outputs, otherwise that thing would exist independently.
Why isn’t ranking the score of an llm output against what it is “supposedly” predicting?