> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
The output token limit and the context window are separate constraints. The context window is how much the model can see at once. The output limit is how much it can generate in a single API call.
In an agentic loop, each API call gets its own output budget. A 'turn' is one response from the model, whether that response contains a tool call, a reasoning step, or a final answer. So with a 1M context and a 64K output limit, the agent can run many turns where the context grows each round (accumulating tool results, prior thoughts, user messages), but each individual response is still capped at 64K tokens.
Expanding the output limit to 1M matters most for tasks that produce a lot in one shot, like writing a full document or a very long file. For most agentic workflows that naturally break into short turns, the per-call limit was rarely the bottleneck. The context window filling up was.
Hmm, but I thought that each token generated effectively becomes a part of the context window for the next token. So 1mm context + 1mm output means that the 1 millionth output token will effectively have been generated with ~2mm tokens of context. But maybe that’s wrong.