logoalt Hacker News

satvikpendemtoday at 5:14 PM2 repliesview on HN

Reduce or turn off thinking:

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates


Replies

Casteiltoday at 5:44 PM

Given that it apparently defaults to 'xhigh', this is probably the answer.

Granted, it's still much lower tokens/s than you'll get out of many MoE models.

Edit: Even set to medium or low there's still a lot of second guessing, less consistency, lower 'acceptable response' rate, and slower/more token churn vs gemma4:26b-a3b. I think gemma4 is just a better 'general purpose' model.

IronWolvetoday at 5:41 PM

Thank you, this is exactly what I needed.