logoalt Hacker News

astrange • yesterday at 3:59 AM • 1 reply • view on HN

No, most of a modern LLM's training time is spent in RLVR, which does not "acquire information from an existing source". You can RL behaviors into a randomly initialized neural network.


Replies

hodgehog11 • yesterday at 4:01 AM

This is true, but you're not going to get anywhere. The pretraining phase is necessary to immensely reduce variance in the RLVR stage. Once there, RLVR has a surprising tendency to only restrict the generated space further. This is not true of RLHF, by the way, which I find to be particularly fascinating, but I digress.