I mean the thinking does not bloat the context window because it gets dropped at the next request.
It doesn't. It's called preserved reasoning and every recent reasoning model does it
It doesn't. It's called preserved reasoning and every recent reasoning model does it