logoalt Hacker News

sebastienburel • today at 5:42 AM • 2 replies • view on HN

On a Mac the baseline I'd want is MLX, not llama.cpp. llama.cpp isn't the fast path on Apple Silicon for most models people run locally, so a speedup over llama.cpp could still be slower than mlx_lm. Do you have that number?

Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?

And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.


Replies

bythreads • today at 6:51 AM

i think this is a wash at best, but might be ok if you have a very specific model that is completely unoptimized - but these days you could just point your astra level llm at it and say "make this faster"

anerli • today at 8:14 AM

Yes MLX is generally a better comparison point overall for Apple, planning on releasing a benchmark for that soon. However against the MLX-based engines we've compared with so far Magnitude will continue to have an edge, especially for decode kernels.

Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.

The endpoint is standard OpenAI compatible chat completions.