I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct. I don't know if a lot of people really want to be advertising "oh god our bill is horrible")
I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.
With their models being open-weight, any two-bit firm in the world with enough capital to invest in a few servers can become a provider capable of carving out their own little niche in the economy.
The Chinese see this as a lift on the entire economy, as it comodotizes the technology to a degree in which many firms can serve many sectors of the economy, a true total-economic win worth the public investment.
The American strategy is built off of private investors believing that with enough money poured into as few companies as possible, one or two firms can come to dominate the entire market and start charging an ever burdensome "tax" on every sector it can touch. Not what I would call a total-economic win for the country.
China is the factory of the world. They don't need software to win. Rather they prefer software is free and they can win in hardware. So if AI inference is free, they can put it in as many hardware components as possible and sell them in the market - think toys, cars, tools with chips manufactured in china optimized for the use case. In long term you tend to commoditize hardware. We have thousands of device types of cheap x86, and with linux/bsd software on it coming from china. Why do we think GPUs will be different.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.
In China, it's because they are being heavily subsidized to do the research activity. It's not really complicated -- if you allocate public money for people do to a thing, they will do it.
Undercutting US dominance in AI is huge for China. If the entire narrative is that you have to use Anthropic or OpenAI to access a decent model, then China's AI labs are sitting on the sidelines as some third rate solutions. China publishing the model weights of models comparable to the frontier proprietary models drastically undercuts closed labs dominance. Maybe these Chinese AI labs don't have the billions infrastructures some of the US players do, but they don't have to if the model is open weight. Many inference provider companies around the world have hardware that can run these models and they will happily run frontier class models for people. Starting in 7 days, people will have the option of which of many providers they want to use to access K3.
Making frontier grade models a commodity will make a competitive market where companies compete for business by improving their quality and decreasing their prices. The cost to access frontier grade models will continue be driven down the more competition that enters the market. This commoditization will challenge the valuations of Anthropic and OpenAI.
The Chinese domestic market is extremely competitive on most things, including LLMs. Once one top company went open weights there, that's going to put pressure on others to follow suit. It'd be akin to if a company like Anthropic or OpenAI went open-weight, it'd probably result in a domino-effect of more US models going open-weights since otherwise that competitor is going to win a huge chunk of mind/market share for free.
It's also relatively free right now. Few people are going to run local models, and in the future it's likely that every model being released today will be obsolete. The only real downside is ease of distillation for competitors, but that's probably impossible to stop anyhow.
They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a guaranteed-correct implementation. Similar business model as open-source SaaS companies.
Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though
if you do the training then you're in control of the output. For example, recommending your products/services or failing to mention your competitors. You could also automatically introduce backdoors into code deemed interesting, i'm sure all governments are very interested in having that influence.
> I just don't really understand the business model behind it
There’s a lot of value in the same sense there is a lot of value in controlling what Google search results are shown and what people see in the Twitter feed.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though
I agree. The thesis in the article is interesting insomuch as I had not heard it expressed this way before: US restrictions on GPU exports have made it feasible to train models in China but not serve them. Therefore open model is a hack to get around the export restrictions, since models can be trained internally but shipped out of the country to be served elsewhere under the banner of open weights. I don't really buy this argument - inference is much cheaper than training and they are hosting their models anyway.
I think it is more likely (a) they have the money to do it and they need it for internal reasons - these are huge companies (b) there is a lot of prestige in China associated with besting American technology (c) people are still basing logic on outdated ideas of Chinese capability which are no longer true.
So it is easier than people think for Chinese labs to do this, they need to do it anyway and there is a lot of prestige from opening the weights. It is honestly not that different to why American companies themselves have released open weight models.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though.
Why is it so baffling that people want to build great things? There are plenty of people who are happy building things for a salary and have no interest in taking over the world. Do you find the whole world of open source software baffling? Linus Torvalds and Richard Hipp and Antirez created the world’s most prolific software products and released it for free.
I don’t think they need the hardware to become affordable (as a regular end user).
They need their models to be good enough and cheap enough. Then the rest will follow. Companies will figure out how to host them for you efficiently, and you pay them monthly.
I don’t think I will ever want to set up a home server, no matter how inexpensive the hardware gets. At work I still use Cursor (with Anthropic models usually) because it’s paid by my employer, but for private stuff, I’m already using cheap models with OpenCode, and it’s extremely cheap and surprisingly capable.
I think there was around half a year between where the best models became good enough (last year December?) and where the cheap models became good enough (couple of months ago?).
I've always found it weird how ludicrously poweful personal hardware has gotten. The fact that you can't buy a midrange CPU with less than 16 cores, or that a high-end gaming GPU is 100+ TFLOPS just blows my mind. The fastest supercomputer in 2004 was 70 TFLOPS. Absolute crazypants level of power, and companies were fall over each other to get people to buy it.
'Libre Office' did not 'win'.
People are happy to pay $50/year per seat to have the extra features and to not have to deal with stuff.
There's an issue at the margins here:
$1000/employee is a massive cost - it has to be deeply justified. $50/employee is like ... $2 out of your pocket. It's an incremental cost. The CFO is happy to pay it if there is a lot of value.
A lot of software is in that later category.
Imagine if gasoline was 1 cent per litre - and there was 'free gas' but it was a pain to use, and you had to check a bunch of things. You may just pay the 1 cent.
AI is not quite that yet, but these dynamics will play out eventually, for a lot of things.
My guess it's to mess with the proprietary vendors economically. It's in China's interest to squeeze the US AI business.
The business model is making your country more innovative, which makes its citizens richer
Americans used to do that too, spectacularly.
Just consider the two alternatives: one is your industrial sector with this incredible new automation and analysis tool available for free. The other is one where it has to pay huge chunks of its resources to overseas companies.
Would anyone besides those within China themselves use a completely closed and hidden model from China for their critical business needs?
I think part of it is definitely to weaken US providers and the US economy as a whole. China has a completely different domestic economic structure and motivations from western cultures... it's probably closest to a fascist economy mixed with a Maoist cultural ideology behind it. There's definitely winners and losers and the state tends to have tight controls over everything though.
I also think the restrictions on OpenAI and Anthropic are somewhat short sighted. In that the guardrails dramatically limit efforts towards securing your own software in many ways. Yes, it's also "dangerous" and maybe there should be a means of identifying "domestic" or otherwise "secure" accounts for those allowed to use the models without the same guardrails in place.
> Is it just a direct play to undercut the US providers because they view them as a threat?
It's China putting pressure on a financing strategy in the US that was always a house of cards. It could also be China democratizing something that should have ALWAYS been democratized. Maybe both when the history books on this get written and absorbed by the winners of the LLM wars.
i'm building thigns with open models: 128GB AMD 395+; expensive 72GB blackwell; old NVIDIA 48+48GB; cards.
If it weren't for the massive memory cartel of OpenAI/Anthropic et al, both Mac and AMD would be selling these things.
I repeat, the models are building, modifying and deploying almost anything on github.
Chinese companies are cut out of the market by the US gov restrictions, so making the models open is a survival tactic.
If you're NVIDIA then open-weight models are a classic example of commoditizing your complements; cheaper models mean more people buying GPU's to run them. [1]
Your guess is as good as mine for China though.
[1] https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/