The model is finetuned to enter/leave its thinking mode using special token separators. there's no reason to assume the tool calls induce the same token distribution or produce the model's actual native reasoning trace
We have its actual reasoning traces, and we have these psudotraces, distribution / nativeness is testable now
We have its actual reasoning traces, and we have these psudotraces, distribution / nativeness is testable now