How can the overall trajectory length be the same across reasoning efforts? I don't see how this is possible even if reasoning is not included in the trajectory length calculation.
I think Tibo was just keeping all else fixed and it’s an illustrative example rather than a perfect real-world trajectory.
I think Tibo was just keeping all else fixed and it’s an illustrative example rather than a perfect real-world trajectory.