B, 2-3 agents. The KV cache framing is the right one. Single-stream tok/s is what shows up in benchmarks but it's not what actually hurts when agents are sleeping between tool calls and waking up needing their full context. The question I'd want answered about an engine like this is how it handles partially-cold contexts, because agent sessions aren't uniform sustained reads, they're bursty and interleaved.