logoalt Hacker News

bmitctoday at 7:23 AM3 repliesview on HN

Why is token efficiency a concern with free models?


Replies

dbspintoday at 9:10 AM

They're not free to run, Kimi K3 needs to be run on the cloud, and the quantised versions aren't as capable. Unless you happen to have 3 - 5 TB of VRAM and an 8-node cluster of 8× NVIDIA H100s to run the full fat version. Plus the weights are not yet available to download in any case.

show 2 replies
antilopertoday at 8:51 AM

Because you're paying for tokens. Especially output tokens.

cdud3today at 7:38 AM

Because that's the only argument left after Kimi beats Fabel in results, price and autonomy.