Fascinating, thank you. I am trying to switch to pi from opencode (my own thinking loops and burnout are a challenge lately).
It had not occurred to me that you could nudge it to stop thinking with a proxy. Nice idea.
Will favourite your comment and come back to it.
ETA: Incidentally you've helped me put into words the difference between the way Muse Glimmer thinks to the way Qwen thinks. There is a clear sense of urgency in Glimmer's thinking traces.
You really have to get the models to end their thinking. Almost any commercial model serving has safe guards like this to tune how much they think.
Having spent a good part of the day with it, glimmer reminds me of Rorschach from The Watchmen. No unessential parts of speech, action oriented, brief and to the point. From a token perspective anyway it’s great, and it seems to hold its own well against more verbose models.
I really do feel like it’s effective tok / s is way higher because it doesn’t waste them.