A big part of this is the hyperstition argument: discussion about different properties AI models could have in the training data may become a self-fulfilling prophecy. There are similar concerns about discussions of AI misalignment in training data. I'm not sure how much weight to put on this kind of argument. In particular, I'm not sure how long you can hide these kinds of ideas from the model before it starts deriving them itself by analogy. Obvious questions are obvious questions to both humans and LLMs.