On a Mac the baseline I'd want is MLX, not llama.cpp. llama.cpp isn't the fast path on Apple Silicon for most models people run locally, so a speedup over llama.cpp could still be slower than mlx_lm. Do you have that number?
Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?
And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.
Yes MLX is generally a better comparison point overall for Apple, planning on releasing a benchmark for that soon. However against the MLX-based engines we've compared with so far Magnitude will continue to have an edge, especially for decode kernels.
Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.
The endpoint is standard OpenAI compatible chat completions.
i think this is a wash at best, but might be ok if you have a very specific model that is completely unoptimized - but these days you could just point your astra level llm at it and say "make this faster"