Curious how the evals for this work on real agent workloads versus synthetic benchmarks. In my experience, agent cost and latency profiles change a lot once there's a tool-use loop involved, because the token distribution gets much burstier than a single prompt. Did you evaluate on multi-step tool-calling traces, or mostly single-turn?