You see the same in GLM models. Even slight ambiguity in user instructions will send it into a tailspin on what intention was in thinking tokens then it goes let’s just make a judgement call on a direction and then proceed