logoalt Hacker News

retatopyesterday at 11:20 PM1 replyview on HN

But wouldn't higher tps allow for more reasoning or other hidden processes, potententially making a smarter model?


Replies

dabbzyesterday at 11:58 PM

This is my thought as well. Models have to be intentional about which tokens they burn because there's a real lag time. If you can just fork out 10 different reasoning sessions at once with no regard for token waste/lag, you can compensate a smaller model with just doing more at once with it. No idea if this is reasonably true though.