Training to ignore evidence and logic in one domain transfers to reasoning degradation in other domains.
You're assuming that your prompt is not being intercepted and rerouted by a lightweight prompt classification model.
In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.
Is this actually documented?
Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?
You assume your highly charged political query is hitting the main LLM at all and not some external short circuit.