Qwen thinking is really good in Mandarin; and probably natively trained the most there.
Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language.
Is the quality of the thinking better or it's just shorter since Mandarin is more compact?
This is most likely because the vast majority of the information the model absorbed during training was in Chinese. As a native Mandarin speaker, I frequently need to convert the prompt into English and output it in English in order to avoid that the model falls back into Chinese reasoning logic.
PS: Switching the thinking process from Chinese to English can also significantly circumvent certain self-censorship mechanisms built into the model.