Maybe detecting watermarked text is a skill to be attained. Not allowing proper feedback and training will not allow people to notice the difference on time, thus ruining the data?
Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.
That's fine. But the author's goal here is to determine if anyone can tell. So he's probably not going to ruin his experiment.
That would be a different experiment.