logoalt Hacker News

mirekrusinyesterday at 6:37 PM2 repliesview on HN

Personally I find speculative decoding much better strategy than MoE – performance wise it's there at 90-100 t/s on 2x4090, great intelligence – really great fit.


Replies

d4rkp4tternyesterday at 6:42 PM

A lot of people, including me, don’t want to bother with GPUs, they’d rather run it on their M1-M5 MacBook. For example the 35B-A3B is very usable even on a M1 64GB MacBook.

show 1 reply
c0m47053yesterday at 7:41 PM

MoE is great on systems that lack the VRAM to host the full model. On my 16GB VRAM system, I can get 100 tok/s with Q4 Qwen 3.6 35b a3b, and 15 tok/s with 27b.

MTP is a trade-off, as it pushes some more of the model off the GPU.

I have managed to get usable quants of Laguna S2 and even DeepSeek V4 flash on this setup.

There is clearly some intelligence loss compared to similar sized dense models, but I feel like it stomps on the 9-12b models I could run fully on GPU