logoalt Hacker News

hugmynutusyesterday at 3:14 PM1 replyview on HN

Qwen3.8/Qwen3.6 has a weird self doubt/thinking too much problem. You can prompt it away. I would say it "approximates" Opus 4.X class models well enough especially for coding/linux problems.

The only reason I stopped using it as much is I was getting 25-35tok/s on Intel B70 (non-quant) which made some responses slow. For a long running/autonomous task, it would probably be sufficient.


Replies

eightysixfouryesterday at 6:41 PM

Two things:

- check your temp settings vs. Qwen's recommendations, they specify what it should be in the model card for thinking on, and that reduces some "over" think.

- the model appears to be intentionally designed to do a lot more test-time compute, if you anthropomorphize the tokens, it looks like overthinking and anxiety, but it is just spending compute to get to the end result, so it may not actually help it to prompt down the token spend (depending on the problem)