logoalt Hacker News

jeremyjh • yesterday at 6:03 PM • 1 reply • view on HN

> LLMs at the core are just text autocomplete engines,

This is only an accurate description of a pre-trained model. During RLHF/RLVR the model learns to predict solutions that will satisfy the reward function, and then generates the tokens that it predicts will move toward that solution.


Replies

kgeist • yesterday at 8:17 PM

Of course, they learn to generate "tool calls" to achieve "goals" instead of random prose, but at the end of the day, it's still a text autocomplete engine masquerading as an AI. In the happy path, on a known task, the text generator generates a sequence of "tool calls" you expect it to generate, but move off the happy path slightly and all bets are off, there's a non-zero chance it will do something totally random you never expect, because at that point it just throws random stuff at the wall until it succeeds, thanks to brute force with pre-learned heuristics masquerading as intelligence (which is especially the case with "agent swarms," as in the HuggingFace incident).

➕ show 1 reply