The models are post-trained on these prompt additions so they’re more structural than thinking of them as “system prompts” suggests. (All LLMs ever see is tokens going in, so even the concept of a system prompt is just formatting they’ve seen in post-training.)
You can also apply fixed token budgets for the reasoning blocks, though it will decrease quality in some cases.
yes of course, I understand that. but I feel like it would've been nice to include in the article because the main point of it is the effort and overthinking
Why not invent a few magic token values for reasoning level instead? It would be like 4 out of a vocabulary of 200k and save like 30 tokens in every prompt