Yes MLX is generally a better comparison point overall for Apple, planning on releasing a benchmark for that soon. However against the MLX-based engines we've compared with so far Magnitude will continue to have an edge, especially for decode kernels.
Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.
The endpoint is standard OpenAI compatible chat completions.