I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
Did you try lower reasoning modes as well?
> it spends way too many tokens on overthinking stuff
Yeah. It’s a German model.