You can imagine that as people get used to working with Claude, they defer to its judgement. So the people choosing which RL path is better may say "yes, Claude, that was a good refactor!" because it did something hard that it may have been able to superficially justify. Actually the change was unnecessary and complicating.
The Claude trainers, as they themselves adapt to Claude's output, are collapsing in their own distribution, so even "new" from-human data is already contaminated.