We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals).
We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.
We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.
And no, no AI company has ever come to us and asked to run training on all of our scanned copies.
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.
The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd.
Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.
What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.
I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".
Maybe we can kill two birds with one stone: digitize rare books and reverse the damage from Authors Guild v. Google [1].
Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Can only blame AI companies to a limited extent. This is apparently the legal way to do things because of stupid copyright laws.
Support your local shadow library: https://annas-archive.pk/donate
It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)
From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.
I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
Reminds me of Blood Meridian where The Judge meticulously sketches the rock glyphs that he comes across, and then destroys the original.
Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."
They should be forced to publicly release the books as an Ebook.
Leave it up for anyone to download and then compensate the copyright holders later.
In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.
Isn't this the legal requirement for digitizing books usually? I dont like that its being done, but I feel the direction of anger for this one is misplaced. Or at least partially misplaced.
Which book that was rare was destroyed? I'm interested to know a few titles.
>And the judge said it's legal.
Was was the alternative? Order them to stop quietly destroying their property, it is making someone else very upset?
Librarians were already doing this at scale in a process euphemistically called "weeding":
https://www.ala.org/tools/challengesupport/selectionpolicyto...
Is there any proof of this at all beyond this random message?
Its ironic how the whole systems reacts when someone like open library tries to actually do digital preservation
This is the opposite of a book burning. These books which only a few would ever know the names of, let alone find, let alone read, are being digitised so they can be found in electronic searches.
Also discussed here https://news.ycombinator.com/item?id=44381838 from June 2025.
This was known a long time ago? https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
This is the result of our copyright law in the United States, which is extremely tilted to favor authors and publishers. The judge made exactly the right call and the companies are following the law.
The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.
I dunno.
It's sad in a romantic kinda way, because of the lost artifact, but the information is what makes the book valuable, not really the medium.
The out of copyright books don't really need to be destroyed anyway for them to be fair use for AI training, and arguable, even if you needed to, you only need one copy per title per company at most.
So it's not a gigantic loss.
Where's the evidence the books are "rare"?
I had some old computer/unix/etc books for which I couldn't find digital copies. I was considering paying to have them scanned because lugging them around each time I moved was getting to be a pain in the butt. When I found out they destroyed the book in the process I could just never go through with it. I later found out there are non-destructive scanning machines (even open source ones!) but ended up selling the books before I ever went that route.
How much of this shredding them isn’t just copyright but rather they don’t want anyone else having this information in their datasets?
Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.
As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses
You can’t champion on the greed that is copyright and then be sour when Anthropic tries to work within this framework.
There is a word for that kind of behavior.
At the very least why not upload the scanned books to the internet archive while already at it?
This is highly disturbing news; is this standard practice? What did Google Books do before?
From 1984: "Every record has been destroyed or falsified, every book rewritten, every picture has been repainted, every statue and street building has been renamed, every date has been altered. And the process is continuing day by day and minute by minute. History has stopped. Nothing exists except an endless present in which the Party is always right."
The source for this is "I've heard it from some guy". Before we freak out, perhaps make sure it's actually happening?
Here's the original story from 404 Media that this tweet is paraphrasing (without linking to, of course, because X disincentivizes links): https://www.404media.co/ai-companies-are-buying-tons-of-old-...
The segment that talks about rare books:
> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]
> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
It's just modern book burning. The contents don't even matter, you will never see the text and images in these books again.
It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience:
> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.
Article is paywalled, but I saved a few quotes here:
Digitizing books, even if it means the destruction of the original, means more people end up having access to the knowledge within, and is a good thing. Full stop.
Is this true? Rare books would very often be out of copyright for a start. What the the actual ruling that says you can scan if you destroy the original? ISBNdb is a database of book information, as you would guess from the name.
Could someone explain to me, why rare books are so precious for them?
Say i feed the largest LLM a book of an alien civilisation, that it definitely hasn't seen before. Then this tiny piece of text muds the vast ocean (latent space) of the model minimally. It will not be able to cite from that book reliably after that fine tuning. Especially for rare books, because them being rare implies, that there aren't 1000s of other books, that encode the same information.
It's general language modelling capabilities might get an iota better, of course. But for putting factual information into it, wouldn't RAG be a much more solid approach?
The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit:
> It is equivalent to book burning in the past. A form of thought control
Literally Rainbows End. Vinge was an oracle.
It's a real shame that no one ever got that book in front of Hayao Miyazaki's eyes.
I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?
I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them (without destroying them) before these things come for them.
I wish...
There needs to be wider reporting of this if it is true. Where are the journalists ? If this is true, it is scorched earth on the world's knowledge base. And the business models are not proven yet. I can now see why some ai companies paints a dire picture of the future - they are basically annihilating all knowledge sources at the altar of their ai gods. I can also see why there is growing ai-skepticism.
Destoying books, for any reason, is a crime and it should not be done, ever!
> bulk-buying rare books
Isn't that oxymoronic? If they can be bulk-bought they aren't rare.
This story was reported by 404 media. OP linked some random person’s tweet.
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
May be I am old schools but rare books must remain rare and LLMs shouldnt have access to those books
Maybe we should not have forced AI companies to shred rare books so they can use them for training?
And this is the beginning of the end for content creators.
Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?
If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.
Edging a little closer to the Krazam video "rare data hunters."
Are the AI companies also destroying the digital scans they made?
I've limited sympathy for the publishers.
It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.
And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
If we have to have copyright laws, I'd like to see two changes to them.
When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.
And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.