Reasoning models will frequently backtrack and re-assess what they've said so far. That's one reason test-time scaling is so powerful.