Where is the actual evidence of distillation? I keep seeing this repeated ad nauseam but I must have somehow missed the evidence.
Distillation a pretty well documented technique that actually pre-dates LLMs https://arxiv.org/pdf/1503.02531
Here is a project that guides you through it if you want to prove to yourself that it works https://github.com/arcee-ai/DistillKit
It turns out you can train a 1b model at almost 1000 tokens/s on a m5 max laptop. As a personal experiment, I've been asking Sol for synthetic training data and synthetic agentic training data (model distillation in it's purest form), plus modified opencode, codex transcripts etc for training data, and nobody's even paying me to do it. If I'm doing it has a hobby, you can bet industrial users are doing it.
The evidence is Anthropic's own reporting [1]. You may doubt that they're telling the truth, but that's what they're reporting.
[1] https://www.anthropic.com/news/detecting-and-preventing-dist...
Musk confirmed in federal court that xAI does it: https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...
It's also how providers build their smaller models out of their larger ones; they publicly talk about the process.
Here's an example: https://github.com/microsoft/Build25-LAB329
[flagged]
there is no evidence. it shortcuts post training by a huge margin this is true. but that is all.
Been using a lot of Kimi K3 lately and the answers have been… „load-bearing“ to the point of hilariousness. It‘s obvious from where they distilled, even if sceptics rightly point out it can‘t have been the only source of their secret sauce, as it‘s been better than the current Opus 4.x at the time of release.