logoalt Hacker News

xlaynyesterday at 7:34 PM4 repliesview on HN

Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading your announcement. Beyond the space saving, why removing the MTP? improves speed exactly for the group that could benefit from it.


Replies

gruturoyesterday at 9:27 PM

The reason for running those insanely low quants is to fit in extremely limited memory budgets. The first thing you sacrifice is speed, then context and accuracy (up to you in which order). IQ2_XXS and below is desperate/proof of concept territory. If you have a spare half gig for the MTP drafter, run a larger quant instead, it will be less incoherent, and damn the speed, it won't be garbage at least. Only around Q4 I'd allocate the comparative luxury of more memory for a speed increase. At least on a dense model. MTP makes a lot more sense (but helps statistically a bit less) on an MoE.

Qwaiting for that 3.8-35B-A3B

danielhanchenyesterday at 10:26 PM

Hey we did not remove the MTP for sizes above 8GiB - but yes for small GGUFs under 8 ish GiB, we removed the MTP module (IQ2_XXS and lower), because it's 500MiB to 750MiB in size, and on small 8 GiB machines, even 500MiB is needed.

As someone in the comments said we made a separate Q4_0 MTP if that's helpful so you can use that.

But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL

show 1 reply
mike-the-brainyesterday at 7:35 PM

you can still have it, no?

> We also removed the MTP module from smaller quants under UD-Q2_K_XL (8.37GB and lower) to converse around 500MB of disk space - you can use the Q4_0 MTP separate module if needed

show 1 reply
walrus01yesterday at 10:47 PM

Q2 quantization is basically giving a capable model a lobotomy. It will not accurately represent how smart or capable something like qwen 3.8 27B in Q8 will be.

show 1 reply