> I doubt these will have worse protection than software does, which has far better protections than copyright
Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.
Software is protected by the DMCA, patents, licenses, EULAs, all of those aren't there for books. I doubt new laws won't be written for model outputs.
Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply
Distillation isn't a copy. Distillation is more akin to "clean room" implementation.
Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.
By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.
There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.