logoalt Hacker News

m00dytoday at 3:35 PM0 repliesview on HN

I would want to see three things before drawing strong conclusions:

End-to-end tokens/sec and cost on realistic coding agent trajectories, including tool outputs and retries, not isolated decode benchmarks.

Cache hit rates and prefill cost for branching, multi-turn sessions.

Router-load distributions after post-training, where expert collapse or specialization problems often show up.