logoalt Hacker News

DanielHBtoday at 9:36 AM2 repliesview on HN

I am very much a beginner to local LLM stuff and I find it incredibly hard to figure out how to run models optimally with the correct settings for my hardware. The number of different variations of the same model and how each quant work is super confusing as well.

When I tried to run llama.cpp directly I was getting max 9tk/s on qwen3.5-9B, then I tried LM Studio with the same model and got 77tk/s. I haven't figured out yet how to get MTP working properly in either.


Replies

cartoday at 10:54 AM

If you are on Mac, have a look at the Llama-macOS app. They claim sensible settings for the linked model downloads. I'd expect the authors of Llama.cpp and the Huggingface folks to know this stuff.

https://news.ycombinator.com/item?id=49328008

tiahuratoday at 12:44 PM

If your package manager / configurator isn’t claude code or codex, you’re wasting time.

show 2 replies