128k context is not a limit of the model, that's a limit of implementation:
"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."
https://huggingface.co/Qwen/Qwen3.8-27B
We're talking about the Cerebras implementation, which is limited to 128K.
It's in the link.
TPM means Tokens per Minute.
We're talking about the Cerebras implementation, which is limited to 128K.
It's in the link.