Deepseek is, with difference, the most "Western" of Chinese models, so it's a bit perplexing that it was chosen to test this hypothesis.
I didn't run any benchmarks but I played around a little, and after getting around the API-level filter Deepseek V4's answers about "China-sensitive content" aren't any different from what I get from Claude and ChatGPT.
Could just be resources available? Deepseek is the easiest to get up and running on hardware that's pretty readily available:
unsloth/DeepSeek-V4-Flash-GGUF 4bit ~140GB
unsloth/Kimi-K3-GGUF 4bit ~1.5TB
unsloth/GLM-5.2-GGUF 4bit ~400GB
You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data
We found V4 Flash was significantly more censored than the baseline.