logoalt Hacker News

kmike84 • yesterday at 10:18 PM • 2 replies • view on HN

How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:

  Estimated speed on your machine
  Context tokens Tokens / sec
  25 000 17
  50 000 16
  75 000 16
  262 144 12

262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).

Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?


Replies

anerli • yesterday at 10:44 PM

These numbers are just estimates based on your hardware and may differ from actual performance. It's hard to get an accurate measurement until it's actually downloaded and running. They also don't account for gains from speculative decoding. Working on changes to make this more clear.

For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.

francisjp • yesterday at 10:31 PM

To OP: great work on the release! I am generally interested in this kind of optimization work.

Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.

Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.

➕ show 1 reply