logoalt Hacker News

petuyesterday at 7:00 PM2 repliesview on HN

> Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others

No? Gemma 31B and Qwen 27B are about the same speed. Gemma 26B-A4B and Qwen 35B-A3B are about the same speed.


Replies

trouve_searchyesterday at 7:18 PM

What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods).

Output TPS in vllm for instance:

- Gemma4 26B-A4B: 200-300TPS

- Qwen3.6 35B-A3B: 120-180TPS

- Gemma4 31B: 80-120TPS

- Qwen3.6 27B: 60-80TPS

This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.

show 3 replies
stymaaryesterday at 7:04 PM

There's no Qwen3.8-35B-A3B though.

show 1 reply