I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
Amazon also did this for their "Look Inside" feature. To do this they had to spin up massive digital infrastructure——then realized that they could sell that infrastructure and make more money than from the books they were scanning.
This is was what ended up spinning up AWS as a business.
I dislike the idea of destroying rare books--but how rare? A digital copy has a lot more benefit.
The irony of countries blocking Anna's Archive (UK, Italy, Netherlands etc) but then it's the hackers who conserve and steward the books. Because governments can't stop AI companies from shredding history like they did with Google Books _because_ they actually tried to do it _by the book_.
Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space.
It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
> The project was met with significant legal challenges from authors and publishers which was eventually overcome
I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.
I've worked in academic libraries since 2010. I've always felt like it was a mistake for our digital library leadership to put so much trust in Google Books despite their promises to maintain the integrity of libraries (this mostly happened by the way).
At the time it was obvious and innovative but over time it was clear Google was establishing a technical precedent to corrode what libraries have the power to do. I'm hoping we can continue to do the good work but it's exhausting.
If you read the judgement against (I think) openai, the judge said that it was OK to scan the books for LLMs, if they were destroyed afterwards. That is, only one copy of the data existed.
Internet Archive version: https://openlibrary.org/
Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donat...
Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-a...
Web app: https://archive.org/want/?mode=donation_book
For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap.
If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.
Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.
I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.