2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,
On our benchmarks we approach 2x decode speeds on a variety of Mac hardware (tested most on M4 Pro and Max).
llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.
B, 2-3 agents. The KV cache framing is the right one. Single-stream tok/s is what shows up in benchmarks but it's not what actually hurts when agents are sleeping between tool calls and waking up needing their full context. The question I'd want answered about an engine like this is how it handles partially-cold contexts, because agent sessions aren't uniform sustained reads, they're bursty and interleaved.