It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.
Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.
Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.
Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.
Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.
Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.
But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.
You ask "Why destroy physical books?"
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
A few points for people: 1. some books are out of print. 2. some books CANNOT return to print. 3. all books prior to the 21st century are products of human minds. 4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access. 5. not all people value all books equally. 6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them). 7. many books are already lost, and historians are the ones who most voice the harms. 8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it. 9. if no existant copies remain, then the price for access becomes effectively infinite. 10. removal of books denies human agency over access to information. 11. costs will follow a steepening curve much as ram did.
first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.
Does adding an old book risk making the model worse? Off the top of my head:
- Reinforcing outdated, disproved or otherwise incorrect information.
- Reinforcing outdated forms of communication e.g purple prose.
You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory
The post reads like a propaganda. Evoking feeling over metrics or results
(<- that's what propaganda does by definition)
e.g.
> but ethically, it’s an extremely serious crime against humanity.
Who are you that you consider it a serious crime against humanity? What are you a saint?
> Knowledge is permanently monopolized on private servers.
Throughout human history, it's the knowledge/info that provide one with wealth and advantage over others.
With the author's logic, every single private knowledge the author is not sharing is an extremely serious crime against humanity.
The logic of the writing is not correct.
I collect old books. It's not common to buy books at all yet to buy old books. Library sales as big as concerts exist, little book libraries everywhere but the used book stores are constantly closing.
I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.
America is interesting.
* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz
* download and publish a book as a company -> fine 1% of revenue
* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations
* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model
I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order
Also the summary:
> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
This is a good example of immature writing. Buried on the bottom of the page are two links, both revealing internal tickets that are hinting as to how I can actually help. There should be big, bold, easy to follow steps for volunteers.
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?
And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?
Why is that?
If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
I can't recall the company, it's been several years, but they were scanning rare texts in detail and making them available online. It wasn't Project Ocean/Google Books - that said I don't trust Google to be a good steward here.
In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.
I think this is a duplicate of a story that ran on the front page yesterday:
https://news.ycombinator.com/item?id=49383026
The big thing here is: libraries and the book trade destroy millions of books every year. If you clean your attic out, box up all your old books, and bring them to your local library donation box, they'll quickly sort through it for things that might actually circulate, and the rest go right to the recycling center.
A lot of people in these comment threads seem not to understand that destruction is part of the natural lifecycle of a book. Books generally don't get preserved.
And then there's the problem that the original reporting that kicked all this off, at 411, is specific about what is meant by "rare books". It's not, as they say, first editions of Oliver Twist. Rather, these are books nobody cares about; that's what makes them rare in the first place. Vanity press stuff, or manuals for old equipment that isn't produced anymore. All these books would naturally end up a dumpster.
I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.
What is the deal with AI Companies buying old books to scan and then destroy them?
https://old.reddit.com/r/OutOfTheLoop/comments/1vszifd/what_...
At the risk of angering the zeitgeist, destroying one or two or ten paper copies doesn’t make AI companies “become the only ones in the world with digital copies.”.
I think a solution to this could be for the government to require the companies to provide the original scans and OCR results to the government for safe keeping until copyright expires.
Clearly that makes assumptions about the function of government and the complacency of copyright holders.
A more honest framing would be: "The judge ordered the destruction during scanning because of stupid copyright laws"
Does the demonization of destruction of books mean I shouldn't delete downloaded e-books off my phone's Kindle app? All those precious bits, gone forever...
Is there any evidence the books are truly "destroyed" & not merely "disassembled"?
It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.
> A guest post by Anna’s Archive volunteer “u” (translated from Chinese).
I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?
Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.
The price they should have to pay for destroying a rare book (just to scan it) is making a pristine high resolution digital copy of it available to the public at no cost whatsoever.
I don't see any problem and don't see destroying books.
I think books will exist but they will be written with the help of AI.
The context will still be a human mind behind the words in the book.
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.
I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.
Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.
Destroy human knowledge, skyrocket RAM prices, push up residential electricity prices, potentially infantilize a generation of young people. It's all worth it, of course. Every time you _can_ invent a technology, you _must_ invent it. No externality is worth considering.
Looking at my modest shelf of weird midcentury travel logs, I just want to be spared from this moralizing. This is just another case of everyone seeing store shelves as an extension of themselves, just like the decline of other physical media. If their knowledge was so precious then why was it not on the used bookstores and libraries to save them, especially when scanning has become so accessible in the last decade? Why did I find some of these books rotting in overpacked shelves and boxes?
Do you think they will make all of that publicly available after it's been scanned?
This reads as bad propaganda. Surprised that they left the "translated from Chinese" disclaimer. There is the false implication here that the AI companies are somehow tracking down all copies of a particular book and destroying it. AI has spawned a global freak-out.
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
It all sounds like some kind of conspiracy theory, but over the past five years, a lot of conspiracy theories have turned out to be true. That’s pretty creepy.
What’s really disgusting is how unnecessary this is.
LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.
And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.
again "destroy" is not accurate. to scan massive amounts of books you have to cut the binding part as feed scanning has been around for ages
there are non destructive scanning options but they are nowhere as fast and prone to errors
I hate that to save this article about an AI company acting abhorently I have to 'favourite' it.
This is only because this is a legal right granted as part of the purchase of copyrighted material, and because we have tried to stop AI companies from doing this with purely digital copies.
The insanity of attempting to prevent AI learning (which is a direct consequence of the nature of observable information) because of the myth of intellectual property is the main driver of this type of behavior.
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...