This is a pretty embarrassing showing for Grok. I wouldn't trust xAI models as far as I can throw them, but I am interested in how much of this is deficiencies of the model and how much is their harness just terrible. Not that it makes it better, a good harness is far easier and less expensive to make than a model.
Some models have been in long-horizon, multi-day LLM gyms for a long time, and across millions of sessions (if not in the tens, or hundreds of millions now). Some models have not.
The former will perform well in these long horizon benchmarks. The latter won’t.