Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.
Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.
Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.
Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.
Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.
Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?
From what I understand, to work with copyrighted books they need to essentially format shift (i.e., scan and destroy the physical book). So a book vault would not solve this issue.
A book vault would still be useful for out-of-copyright works, but this would only cover a (probably relatively small) portion. Also, I'm not sure how easy it is to reliably determine copyright at scale, so they might just decide that it's not worth it.
At this point my only hope is that in the long run these scans make it to the public somehow (leaks, copyright changes/expiration, whatever), where they can then be accessed and preserved by everybody. Then we could have our true digital library of Alexandria.
> establishing goodwill with the public
Not sure even rare books will dig these big AI companies out of the hole they're digging for themselves.