Two things:
- check your temp settings vs. Qwen's recommendations, they specify what it should be in the model card for thinking on, and that reduces some "over" think.
- the model appears to be intentionally designed to do a lot more test-time compute, if you anthropomorphize the tokens, it looks like overthinking and anxiety, but it is just spending compute to get to the end result, so it may not actually help it to prompt down the token spend (depending on the problem)