I think that's relatively emergent too though! BERT never really did that (at least to my recollection), presumably because its training was never sufficient for it to develop corrective reasoning in a chain of thought.
BERT isn't a next token predictor. It predicts a single token based on the whole surrounding context in both directions.
What exactly do you mean with "emergent" here?