Yea, this reads as LLMs are a pretty obvious technology to develop(for the highly intelligent researchers who are there). Also there's probably a lot of actual divergence in model capabilities and skills that concealed by the fairly narrow set of tests we run them against nowadays. Like wasn't Grok 4.20 super targeted at non-coding tasks.
Why is everyone ignoring the pattern that has existed since training models became a thing? At first it sucks. Then it's better than humans. Just by using it you generate training data that makes it better over time.