logoalt Hacker News

anon373839last Saturday at 9:46 PM4 repliesview on HN

I strongly agree with the premise that distillation is not an “attack”.

But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena.

API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “cold start” problem faster. By far, what matters more is the quality and variety of RL environments the model learns from.


Replies

tristanjlast Sunday at 3:48 AM

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870

This behavior is exactly what you'd expect from a model distilled from Claude.

There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...

This analysis observed K3 identifies itself as Claude approximately 15% of the time.

K3 reproduces Claude's correct current model id, which the real Claude models themselves do not emit. This suggests K3 was trained on Claude data labeled with deployment metadata (API logs, tagged synthetic data), rather than Claude's chat outputs.

And there's an entire Reddit thread discussing Kimi's similarities with Claude https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim...

This analysis shows K3 and Opus/Fable have unexpected correlated outputs https://typebulb.com/u/lab/you-re-relatively-right/full

show 8 replies
RazorBucksICOlast Sunday at 5:51 PM

So if you were to imagine the second largest economy in the world, governed by a monolithic politburo, with a long history of barely concealed programs of corporate and academic espionage, what do you think their approach would be to American technology worth trillions? Now take it a step further, this hypothetical government sees itself as the civilization representation of an ethnic group, and regularly attempts to monitor and police that group regardless of where they are in the world. Also, that ethnic group, separated from the motherland by one generation or less, makes up the majority of research labor in American labs.

Don’t you think there’s maybe a teeny, tiny chance that their approach is a little more sophisticated than just buying retail subscriptions? Maybe the trade secrets are being exfiltrated directly from the American labs?

show 1 reply
spaceman_2020last Sunday at 2:40 PM

The entire debate about distilled vs not distilled is just academic. As an end user, I don’t give a crap how the model was trained as long as it does the job affordably enough

I see all of AI as theft anyway so it makes no ontological difference if the theft was from a human or from another AI

lerchmolast Sunday at 3:06 PM

If your business leaks to your customers… if you can ask your product for its inner secret sauce and it readily gives away the goose. That is not a moat.