There are low stakes use cases where this kind of stuff just doesn’t matter. Not every use case for an LLM involves sensitive or even non public data.
Eg. I have a need to search transcripts of published recordings to extract entities for tagging purposes, find semantic shifts for chapters and other things. The underlying content is already published. If they want to train on my prompts, that was something they could have done with no issue and minimal effort anyway.
Sometimes you don’t need to care why the steak is free.
I've found capable models are incredibly useful for fixing up old ebooks. The kind that were text documents OCRd off a paperback and then dumped in word and bodged into an epub
The ones that have all sorts of ocr artifacts, weird capitalization, and virtually no css
Ran a bunch of older sci-fi through some earlier and it fixes them up very well