logoalt Hacker News

jwrtoday at 8:15 AM1 replyview on HN

I would suggest careful benchmarking. I actually tested and benchmarked, and the new Qwen3.8-27B model is actually slower with MTP on my M4 Max. MTP only gains anything when generating long code sequences, which is very unlikely as the model spends most of its time thinking, not generating code, even if you use it for coding (which I don't).

I get 20 tokens/s on an M4 Max (larger GPU).


Replies

smcleodtoday at 9:14 AM

That shouldn't be the case, it sounds like you've got something else going on with your setup. Here's my benchmarks: https://omlx.ai/my/fadc2127d384283f5df1fcc2c093a9f95700c6a52... which are inline with the communities: https://omlx.ai/benchmarks/performance?sort=tg_tps&order=des...