logoalt Hacker News

kaufmannyesterday at 6:49 AM0 repliesview on HN

I am still trying to get an intuition for the amount of training used for SOTA LLMs. Are there some sources for your speculations?

Does anyone know the ratios of pretrain, posttrain supervised as well as reinforcement learning? (I should probably even distinguish between RLHF and RLVR).

I assume the latter is the main reason for the power of modern models. Is it possible to turn the results of a gym session into trading data?

(Sorry for moving in off topic regions, but I'm interested in that for a long time.)