Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?
That's max output tokens per response limit, separate from context length
It's the output token limit, which has been 128,000 for Claude models for quite a while note
That's max output tokens per response limit, separate from context length