Is this actually documented?
Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?