If we accept the premise that the top Chinese labs are simply distilling and can't compete otherwise: why don't US labs simply do the same thing? Distill their own models and slash their costs by 99% while keeping the same quality output. It should be a piece of cake if even the open labs can figure it out, after all.
One way or another they're getting the same results as proprietary labs, with a fraction of the hardware for a fraction of the cost. OpenAI can't keep raising funding rounds of $100billions to subsidize their compute costs and get results by brute forcing parameter count. And if they're having trouble keeping up with Chinese labs' efficiency, maybe they should stop worrying about distilling and instead hire some of the smart people responsible.
Because distillation-only is a quick performance shortcut that only lets you get to the level of the thing your distilling for the most part or a little bit worse and does not allow you to actually progress past it. It's like only being able to make VHS copies of videos, and maybe do some basic video editing without being able to actually go out with cameras and make new movies.
To actually have something competitive and improved within the next 3 months and not be perpetually behind, you need your own independent model creation process. So to extend the metaphor, a complete movie studio with cameras, actors, staff, sets, budgets, etc. It's the right strategic move to do when you are GPU constrained, which the Chinese labs are, but it won't let you get past it.
A bunch of pedantic people will come out of the wood work citing a bunch of things saying that is not the case because of some detailed mechanics of how model training works and they will get fixated on some of the words I used, but zoom out to the level of what an AI lab is able to produce and this becomes evident.