logoalt Hacker News

prometheus1992yesterday at 5:40 PM3 repliesview on HN

It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat


Replies

Baloogayesterday at 7:05 PM

Now I'm feeling pretty good about getting 10-11 tokens/sec running Qwopus 3.6-35B-A3B Q6_K on an old Mac Pro 2013 (trashcan) with 128GB RAM (DDR3), 12 core Xeon, dual D700s. Arch Linux and llama.cpp.

show 1 reply
trollbridgeyesterday at 9:20 PM

Anything smaller than a 16” runs into serious thermal problems; even an identically equipped 14” just can’t dissipate enough heat.

carloslfuyesterday at 6:38 PM

interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out!