I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.
Deepseek essentially releases instruction manuals in paper form.
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
So boring to see conversations moved over to Chinese models when that’s not even what we’re talking about here. This is about Mistral.
I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
How it should be. Knowledge should not be copyrighted. The world will be a better place with such information democratized
There would be a lot of competition even without DeepSeek. Workers can freely exfiltrate trade secrets without noncompetes in California.
>instruction manuals in paper form
So the most common way to publish manuals?
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.