Everyone trains on your data.
With Chinese providers at least I'm getting a open weight model out of it.
That's very defeatist. Do you have any concrete reason to think the major providers are lying to every one of their business/API customers about not training or storing the data? The business loss of trust would outweigh any benefits of the data.
(And if they freely lie about such things, I don't know why they would bother taking the PR hit when they announced fable had temporary data retention for their abuse prevention)
This idea that the US labs are just directly committing fraud against effectively every major US organization is the most tinfoil hat thing I’ve heard in a long time. Training on excluded data would (eventually) be trivially provable. Forget loss of trust; this would make the labs a defendant to the most legally well-resourced organizations on the planet.