logoalt Hacker News

lin7c • today at 3:23 AM • 0 replies • view on HN

One thing I'd want to see in the evals is per-turn latency across a full agent trajectory, not just end-to-end time. In my experience the workload flips mid-run: early turns are prefill-heavy (big system prompt, tool schemas), late turns are short decodes against a huge KV cache, so a config that's optimal for turn one can be badly wrong by turn thirty. The self-tuning idea is interesting, but I'm curious whether the tuning happens per-request or per-trajectory. With prefix-cached tool schemas the win should compound; without it you're re-solving the same optimization problem every call.