For the web app task I mentioned:
* Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5
* Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3
* Fable: ~14m input (all cached??), 196k output - cost $30
Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code.
Not sure wtf is going on with the Fable stats (a lot tokens, virtuall all of them were cached - I guess heavy system prompt?) but both claude code stats and ccusage tool output match.
> or the web app task I mentioned:
* Kimi K3: (...) cost $6.3
* Fable: (...) cost $30
It's pretty clear that Kimi K3 beats Fable by a long margin.
Would a fairer test not be to use the same harness for all three? I’d suspect the harness to massively affect token use and optimisation