logoalt Hacker News

2001zhaozhaotoday at 1:32 AM0 repliesview on HN

I would love to see models that can think at different rates and also output a thinking scratchpad alongside output text instead of before all output.

Right now models need to rely on less legible compressed CoT to get high intelligence per token/step, but with diffusion they would just need to output more tokens per step instead.