Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
It's also interesting because in humans the existence of "Aha!" moments that are not preceded by or are only loosely related to a chain of thought is taken as the proof of the fundamental mystery and irreproducibility of human intelligence. Now the same argument is made to deny that LLMs actually think. Go figure.
I assume theyre searching the local gradient to see if theres a better descent before proceeding.
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.