logoalt Hacker News

novokyesterday at 7:21 PM6 repliesview on HN

The open models are good because of distillation, which the US labs are actively working against via not revealing CoT ever and now you can see with OpenAI Astra 6 not even having a lot of CoT equivalents being emitted as tokens. Once the anti-distillation stuff is in place the open distillation models will probably start having larger and larger gaps.

If the companies survive the next few years, which they probably will because they represent too much of US economic growth to allow them to fail, this gap will keep on expanding.

Starting from zero without distillation is a lot harder, a lot more expensive and a lot more work. OSS models is what a laggard does to get adoption. China's gov't might keep on sponsoring it as a counter GPU embargo thing, but when gov't get involved, usually the other side gets involved too.

As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.


Replies

nvme0n1p1yesterday at 8:01 PM

If we accept the premise that the top Chinese labs are simply distilling and can't compete otherwise: why don't US labs simply do the same thing? Distill their own models and slash their costs by 99% while keeping the same quality output. It should be a piece of cake if even the open labs can figure it out, after all.

One way or another they're getting the same results as proprietary labs, with a fraction of the hardware for a fraction of the cost. OpenAI can't keep raising funding rounds of $100billions to subsidize their compute costs and get results by brute forcing parameter count. And if they're having trouble keeping up with Chinese labs' efficiency, maybe they should stop worrying about distilling and instead hire some of the smart people responsible.

show 1 reply
cmiles8yesterday at 8:40 PM

If that was the case the major labs would have done that.

They’re burning cash like there’s no tomorrow. They desperately need to show that they have a real business and not just a giant burning pile of cash doing academically interesting things. If they could simply sell models that are 95% as good at 1/10th the price they’d do that. They’re losing the enterprise sector because they’ve not done that.

show 1 reply
janalsncmyesterday at 8:45 PM

Basically your argument relies on two claims, both of which must be true.

1. Competitors to OpenAI and Anthropic are good because of distillation.

2. OpenAI and Anthropic will come up with some methods for preventing distillation in the future.

Both of these are dubious imo. For RLVR tasks like coding in particular, you definitely don’t need continuous distillation to improve, otherwise OpenAI and Anthropic themselves would not be able to improve because there is no better model to distill from.

chvidyesterday at 7:46 PM

Look up how many Chinese are pursuing a computer science degree.

Chinese AI labs do well because they have the best AI people coming out of a huge talent pool.

show 2 replies
upcoming-sesameyesterday at 9:34 PM

Do the western labs not train on the open weight models? Is it a one way distillation?

achronoyesterday at 7:40 PM

Yes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed.

Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?

If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.

show 1 reply