logoalt Hacker News

KronisLVtoday at 8:30 AM1 replyview on HN

> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.

This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.

On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.

Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.

Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.

My overall thoughts (released over some time):

https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...

https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...

https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...

I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.


Replies

re-thctoday at 9:51 AM

> This matches my experience with DeepSeek V4 Pro at Max reasoning,

That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper.

> On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time,

GLM 5.3 bridges this gap.

> I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates.

Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.