logoalt Hacker News

stkdumptoday at 2:15 AM1 replyview on HN

There is a native reasoning effort setting. It defaults to xhigh, I guess to get the best benchmark results, but you can just run it on medium or low instead, or for simple things even disable thinking outright.


Replies

lifepillartoday at 5:22 AM

According to this guy [0], medium is the level that tends to produce way less tokens in agentic workflows ("low" may output less per response, but then the model makes more mistakes, so it needs to iterate more).

[0] https://m.youtube.com/watch?v=z64J6bC16iQ