Some models have been in long-horizon, multi-day LLM gyms for a long time, and across millions of sessions (if not in the tens, or hundreds of millions now). Some models have not.
The former will perform well in these long horizon benchmarks. The latter won’t.
I am still trying to get an intuition for the amount of training used for SOTA LLMs. Are there some sources for your speculations?
Does anyone know the ratios of pretrain, posttrain supervised as well as reinforcement learning? (I should probably even distinguish between RLHF and RLVR).
I assume the latter is the main reason for the power of modern models. Is it possible to turn the results of a gym session into trading data?
(Sorry for moving in off topic regions, but I'm interested in that for a long time.)
What you're saying essentially comes down to "Grok hasn't been trained on long horizon tasks like other frontier models", which...yes. But they've also acquired a company with probably more session data than any other non-frontier lab. They're also selling their model as comparable to other frontier models. Which very clearly isn't the case.