logoalt Hacker News

martianvoid • today at 12:56 PM • 2 replies • view on HN

I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach

Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8


Replies

gizajob • today at 12:57 PM

> it spends way too many tokens on overthinking stuff

Yeah. It’s a German model.

➕ show 3 replies
tzatzikyyy • today at 1:32 PM

Did you try lower reasoning modes as well?