> In fact, one could almost imagine it baked into the assistant/harness
...it already is, via thinking/effort. That wouldn't be possible if a LLM at usable quality wasn't fast enough to allow at least some amount of "thinking" (i.e. hidden text generation).
But your point still stands: we could get massive quality gains by allowing even more thinking by default, if it was fast enough.