logoalt Hacker News

kamranjonyesterday at 1:36 PM6 repliesview on HN

Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.


Replies

vmt-manyesterday at 2:07 PM

are you working in earplugs? :)) even with 128 gigs of ram it must be super noisy.

show 4 replies
toughyesterday at 4:06 PM

I've put Opus to it and it says it will take 3-4h to do the process to the new weights. hoping it works!

ycui7yesterday at 4:07 PM

you don’t need new weight. try vllm-moet from github. it will autogenerate 2-bit plane.

theturtletalksyesterday at 1:58 PM

What kind of tps are you getting?

show 1 reply
knupparyesterday at 3:54 PM

same. been running it in a dgx spark and it slaps