It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.
Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.
Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.
Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.
Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.
Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.
But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.
Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
You ask "Why destroy physical books?"
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
So we don't want companies to buy books and scan them and do whatever they want afterwards, but we also don't want to allow piracy of digital copies (the $1.5B Anthropic settlement)? Bit of a rock and a hard place for them.
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
I collect old books. It's not common to buy books at all yet to buy old books. Library sales as big as concerts exist, little book libraries everywhere but the used book stores are constantly closing.
I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order
Also the summary:
> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
Does adding an old book risk making the model worse? Off the top of my head:
- Reinforcing outdated, disproved or otherwise incorrect information.
- Reinforcing outdated forms of communication e.g purple prose.
You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
What is the deal with AI Companies buying old books to scan and then destroy them?
https://old.reddit.com/r/OutOfTheLoop/comments/1vszifd/what_...
Many countries require a copy of each book published there to be submitted to their national archives/library. The US has had this requirement since 1790 as far as I can tell.
All countries I checked (Germany, France, Canada) seem to have similar requirements. WIPO claims that this is the case in the majority of countries (https://www.wipo.int/documents/d/copyright/docs-en-registrat...).
Thus, most books should not be at risk of getting permanently lost due to these practices.
This is Anna's Archive (ab)using a current controversy (which I think is based around emotional appeals and distorted facts to create outrage over a non-issue) to tongue-in-cheek advertise their open pirate library.
AI companies are supporting the market for books that no one else wanted. The economically illiterate assumption of the author is that "rare" books are good. If they were that good, they would be priced higher.
https://abunner.substack.com/i/210907372/anthropic-is-suppor...
destructive book scanning is pretty standard, it's the only way to efficiently get a good flat scan of the pages without a tremendous amount of human labor to correct distortion on every individual page. i'm curious about what specific rare books are being destroyed and how rare they really are. and at any rate, anyone can go out and buy books and scan them or destroy them or set them on fire or whatever, they bought the book it's their own property. so now physical books are treated as the public commons while intellectual property is privately owned? are we in topsy-turvy world? i feel like there are much more egregious crimes that people ought to be talking about with ai companies, like impoverished black people having their city water supplies poisoned or the massive financial fraud that will destroy the economy or the massive co2 emissions that will destroy the world or just the fact that it's not even artificial intelligence at all and it's not even a technological innovation, it's just a google hack that dumbed down an existing neural network model enough for it to run on lots of nvidia gpus and produce an impressive-enough tech demo to show to gullible investors who don't know what to do with their massive piles of money.
This is a good example of immature writing. Buried on the bottom of the page are two links, both revealing internal tickets that are hinting as to how I can actually help. There should be big, bold, easy to follow steps for volunteers.
A few points for people: 1. some books are out of print. 2. some books CANNOT return to print. 3. all books prior to the 21st century are products of human minds. 4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access. 5. not all people value all books equally. 6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them). 7. many books are already lost, and historians are the ones who most voice the harms. 8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it. 9. if no existant copies remain, then the price for access becomes effectively infinite. 10. removal of books denies human agency over access to information. 11. costs will follow a steepening curve much as ram did.
first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
The post reads like a propaganda. Evoking feeling over metrics or results
(<- that's what propaganda does by definition)
e.g.
> but ethically, it’s an extremely serious crime against humanity.
Who are you that you consider it a serious crime against humanity? What are you a saint?
> Knowledge is permanently monopolized on private servers.
Throughout human history, it's the knowledge/info that provide one with wealth and advantage over others.
With the author's logic, every single private knowledge the author is not sharing is an extremely serious crime against humanity.
The logic of the writing is not correct.
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?
And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?
Why is that?
If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?
I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.
America is interesting.
* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz
* download and publish a book as a company -> fine 1% of revenue
* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations
* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
I can't recall the company, it's been several years, but they were scanning rare texts in detail and making them available online. It wasn't Project Ocean/Google Books - that said I don't trust Google to be a good steward here.
A more honest framing would be: "The judge ordered the destruction during scanning because of stupid copyright laws"
I think a solution to this could be for the government to require the companies to provide the original scans and OCR results to the government for safe keeping until copyright expires.
Clearly that makes assumptions about the function of government and the complacency of copyright holders.
Is there any evidence the books are truly "destroyed" & not merely "disassembled"?
It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.
Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.
Hmm... interesting. previously this was annas-archive.pk, now its on gl domain. What happened?
At the risk of angering the zeitgeist, destroying one or two or ten paper copies doesn’t make AI companies “become the only ones in the world with digital copies.”.
Does the demonization of destruction of books mean I shouldn't delete downloaded e-books off my phone's Kindle app? All those precious bits, gone forever...
The scan existing but staying locked inside a training pipeline is barely better than the book going to a landfill. At least make the raw scans available.
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...