logoalt Hacker News

gerdesjyesterday at 10:29 PM2 repliesview on HN

128k context is not a limit of the model, that's a limit of implementation:

"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."

https://huggingface.co/Qwen/Qwen3.8-27B


Replies

Aurornistoday at 6:25 AM

We're talking about the Cerebras implementation, which is limited to 128K.

It's in the link.

selcukatoday at 2:27 AM

TPM means Tokens per Minute.

show 1 reply