Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a cheaper version of it. There was never any plausible explanation for why this wouldn’t happen. There was never any practical mechanism to prevent someone from saving a conversation and using it to train their own model.
Even if it didn’t happen here, it was still the case that it was going to happen going forward. It was always going to end like this. Invest in the hardware companies, not the model companies.
The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers.
People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patterns, etc. There will be fields where China is a global leader, and Americans and Europeans will have to learn Chinese and move there, or else be stuck in some satellite office of a Chinese company. We’re all in Europe circa 1895 not realizing the behemoth America will become in WWI.
Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape."
Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether it's complete loss of access, or the amount of control you'll have to give up to access them will be ridiculous.
Sort of like the "stealing music is fine" but "lets freak out now that it's producing visual art", in the end the entire thing is a social construct. Whether this is treated as theft or "business as usual" is entirely societal.
Eventually the gap will close, unless there's a major breakthrough that hasn't been made yet.
The court decided that LLMs are a transformative fair use of the data they trained on, and therefore aren’t copyright infringement.
Maybe Kimi is a derivative work as well
Look how hard Anthropic is to even be able scroll back on your conversation, or look at the thinking tokens or subagents. They want to keep everyone coming back to the watering hole but never to learn how to dig a well.
Calling distillation an 'attack' is exactly what I've been describing as "AI Exceptionalism":
It is unfair, they stole the dataset that we stole.
The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable.
These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs.
Setting aside the fact that it doesn’t make any feasible sense to do API distillation, these models are outperforming frontier models on a number of benchmarks, and often times run more efficiently by several orders of magnitude.
We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional.
A “distillation attack” is like a concrete company calling a competitor building a factory with its concrete a “construction attack”
I suspect that distillation attacks may be slightly exaggerated. Most of the training data used during fine-tuning is now synthetic data. You can't just repeat the same stuff twice, therefore another LLM is writing a text book that is explaining a topic in detail, ideally without any gaps in the material.
are there any "open source" efforts to do distillation? Like some place one can submit one's anonymized chat logs? So they can be pooled and used as an open training set (similar to OpenCrawl)
> There was never any plausible explanation for why this wouldn’t happen. There was never any practical mechanism to prevent someone from saving a conversation and using it to train their own model.
At this point it may not even be happening intentionally given the quantity of LLM-generated content that is appearing online and is likely being re-ingested by models.
I wonder, how does distillation deal with unprobed spaces in the knowledge landscape? Is a distilled model worse in some niche area that was not probed? Presumably, this is why frontier labs dont distill their own models internally to release them to the public as a servicable frontier model.
if distillation is the key, why the fuck all other competitors do not release competitive models? and only Chinese can distill this great?! Am I smoking too much?
Pulling on this thread, if the model companies become commoditized and make no money then who is buying the hardware? Seems like it would be the next shoe to drop
or the application layer - which will capture majority of the value.
yeah hardware companies make for nice stories or green numbers on Wall Street - but value will be captured by application layer.
look at history.
assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts.
what is the end game for this strategy?
if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?
Well, there is precedence: Google can scrape the web, but you can't scrape Google. Laws around compiled databases exist for a reason: you can't just copy the phone book if effort has gone into compiling it, it is itself copyrightable
> Distillation “attacks” are not attacks.
Say it louder for the people in the back. All these complaints about "distillation" from frontier labs are bordering on felony contempt of business model at this point. It's great for us. Maybe it's bad for them but nobody other than shareholders really cares.
The optimal outcome for humanity is for oligarchs to spend trillions training a godlike AI, only for the precious weights to just leak. No "distillation" required.
The hand wringing over whether internationally located AI labs are "stealing" output from American ones is the funniest thing in a while.
It's international politics with people talking about AI success as a matter of national strategic advantage and survival. So at best "this was built off our work" mostly tells you that apparently you've got months of advantage when a new model drops before it can be cloned. That's certainly some sort of advantage, sure hope it represents a consistent ability to stay ahead and causes people to redouble their efforts.
Or...of course none of these companies are worth what they say, but the advantage is also not really that great, and a whole lot of people are just really worried about their stock payouts.
[dead]
Thanks for the models guys, sorry for your losses. Once this reality becomes mainstream and undeniable, surely the bubble pops and then what then. Future model development stops? Becomes private? Becomes a public effort?
> Distillation “attacks” are not attacks.
If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it.
So both things can be true: a) People infringe on Anthropics IP and b) what Anthropic did to build their models is legally questionable (or might be ruled illegal, even though I doubt it).
>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models
So why didnt we have these LLMs in 2005?
> There was never any plausible explanation for why this wouldn’t happen.
What a nice post hoc revision of history. Distillation is still an active area of research, that you can distill models as easily as you can it genuinely interesting and absolutely not something that was taken for granted even 12 months ago.
Even 6 months ago this idea that 'using model outputs as training examples' was listed as the reason that all models would fail in the near future due to some spooky circular training catastrophe.
Don't pretend like this was so obvious.
I strongly agree with the premise that distillation is not an “attack”.
But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena.
API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “cold start” problem faster. By far, what matters more is the quality and variety of RL environments the model learns from.