I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money.
The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this is so hard for HN to understand.
It depends on the use case. And most companies (like 90%+) do not have the coffers FAANG has and price does make a big difference.
I use Chinese models, even smaller local ones, for much more than pair programming. If we are talking about deepseek v4 flash, which is basically a frontier model, it is much more capable than the local models I run on my MacBook Pro. The only issue really is finding the right harness.
I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough.
> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.
On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.
Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.
Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.
My overall thoughts (released over some time):
https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...
https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...
https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.
> just a very last-gen way of using agents.
Fable has only been out for a month but somehow everyone is supposed to have moved to a completely different way of working that supposedly only works for Fable and nothing else…