Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.
In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.
GPT 5.6 Luna is an extremely cheap and still very capable model.
A chinese model being in the same ballpark of capability at half the price sounds believable to me.