The older models were less good at inferring intent. The new ones just seem to do a better job.
They're just more subtly wrong when they are
They're just more subtly wrong when they are