My concern is that reasoning could involve some sequential steps that instant models don't.
Not sure if modern models "think" only by outputting <thinking> blocks, or there is a more complex mechanism at play.
It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.
It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.