logoalt Hacker News

anerli • last Wednesday at 11:46 PM • 0 replies • view on HN

On our benchmarks we approach 2x decode speeds on a variety of Mac hardware (tested most on M4 Pro and Max).

llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.