logoalt Hacker News

yogthosyesterday at 6:47 PM3 repliesview on HN

And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening at the same time.


Replies

janalsncmyesterday at 10:10 PM

Probably a little of both.

Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer.

The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions.

Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well.

show 1 reply
bocyesterday at 7:43 PM

Or option three is they are drafting hard off the frontier US models via distillation.

show 1 reply
sdfefcxvyesterday at 7:42 PM

This happened ages ago.

But OAI and Anthropic are trying to cash in ahead of their IPO window. I think that window is pretty much gone now.

show 1 reply