logoalt Hacker News

mark_l_watson • today at 3:15 PM • 2 replies • view on HN

Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.

Progress on running local models has been amazing.


Replies

fsiefken • today at 3:59 PM

Yes, I am running the same on a 64G mc. It's good, but slow at 25 tps on average! I want > 100 tps - but I don't have $5k to spare for an m5 ultra or an nvidia setup.

So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2

https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...

or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp

https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...

https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...

generalizations • today at 3:17 PM

I haven't seen much in the way of benchmarks of those smaller quants. How does it compare to e.g. various generations of Opus?