logoalt Hacker News

anerli • yesterday at 8:14 AM • 0 replies • view on HN

Yes MLX is generally a better comparison point overall for Apple, planning on releasing a benchmark for that soon. However against the MLX-based engines we've compared with so far Magnitude will continue to have an edge, especially for decode kernels.

Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.

The endpoint is standard OpenAI compatible chat completions.