It doesn't really matter what the AI companies want, if they misjudge the market somebody will just start a competitor and take it.
The relevant questions are: 1) Where are the economies of scale in the technology stack? 2) What's "good enough" to consumers, and how does that stack up with the relevant computing power needed? 3) What are the transaction costs along various system boundaries?
I think that the biggest force keeping inference in large centralized services is simply that provides a better product for the average user who doesn't care about local control (and the average user doesn't care about local control; indeed, for most people it's a misfeature, as then they have to administer their own hardware). HN is full of nerds that want to own their own stack; they're willing to put up with a little loss of capability to run Qwen 3.8 locally. But from the folks I know that have tried it vs. Claude vs. Codex vs. Antigravity, the local models are still pretty weak compared to what you can get by paying for a service. Most people will just pay for the service until the performance becomes indistinguishable and the price becomes less.
>It doesn't really matter what the AI companies want, if they misjudge the market somebody will just start a competitor and take it.
The vast majority of people are uninterested in local vs non-local models (not just LLMs but ANNs in general), only the services the models are providing. Even if performant local models can actually be served on device quickly, if the frontier labs don't want to go that way, they simply can and there's little threat the market can be taken from them that way.
If you can serve local models, the 'advantage' is that they can be run on the users device without that incurring any cost to you thus allowing flexibilty in pricing, but if such models can be made then whatever cost would be incurred to labs running on their own servers would be a rounding error and also inconsequential to them. That is to say, they can also compete with you on pricing. So unless someone comes along with something the frontier labs can't figure out or replicate, then it doesn't matter. Open AI might be 'forced' to drum down prices but only in situations where costs have also been drummed down.