logoalt Hacker News

williamse • yesterday at 5:31 AM • 0 replies • view on HN

The output token limit and the context window are separate constraints. The context window is how much the model can see at once. The output limit is how much it can generate in a single API call.

In an agentic loop, each API call gets its own output budget. A 'turn' is one response from the model, whether that response contains a tool call, a reasoning step, or a final answer. So with a 1M context and a 64K output limit, the agent can run many turns where the context grows each round (accumulating tool results, prior thoughts, user messages), but each individual response is still capped at 64K tokens.

Expanding the output limit to 1M matters most for tasks that produce a lot in one shot, like writing a full document or a very long file. For most agentic workflows that naturally break into short turns, the per-call limit was rarely the bottleneck. The context window filling up was.