logoalt Hacker News

XCSmetoday at 3:52 AM1 replyview on HN

My concern is that reasoning could involve some sequential steps that instant models don't.

Not sure if modern models "think" only by outputting <thinking> blocks, or there is a more complex mechanism at play.


Replies

4k0hztoday at 5:30 AM

It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.