Hmm... interesting. previously this was annas-archive.pk, now its on gl domain. What happened?
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
The price they should have to pay for destroying a rare book (just to scan it) is making a pristine high resolution digital copy of it available to the public at no cost whatsoever.
> A guest post by Anna’s Archive volunteer “u” (translated from Chinese).
I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.
I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.
I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.
I'm interested in learning/teaching technologies. Naturally science fiction examples are interesting.
I have found that often LLMs are familiar with the contents of SF books.
But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.
Should be illegal to burn books for this reason. Burning 1 book as a protest, fine. But burning a lot to destroy information? Hell no.
It would be nice if we have some "tracking" e.g. 30% of all known books are scanned. So far all information I searched in the Internet about the progress has been patchy. It's also impossible to understand if exact book was ever digitized or not.
What's infuriating about the book digitization is that you'd think: "Oh, great, this would allow finding a fact or a piece from just about any book ever published..."
OMG, finding a verbatim quote from any given book these days is nearly impossible. Every LLM would state "tis a copyrighted material and I can't share it as is". And search engines now all being AI-driven won't find it either. WTF are you even talking about? I'm not trying to steal the whole plot of the book for my dissertation, I just need the exact quote from the book, just like the author intended, don't give me your "rephrased" adaptation of it. What happens in a few years when every single quote is some misinterpreted shit and nobody even knows what the original quote ever was?
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
This is frustrating. Just to beat the competition and make a few extra bucks, they’re willing to destroy a century’s worth of human knowledge (good or bad).
I'm sure the AI companies will retain scans of the books for training on newer models
I don't see any problem and don't see destroying books.
I think books will exist but they will be written with the help of AI.
The context will still be a human mind behind the words in the book.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
Aren’t publishers required to deposit a copy with the Library of Congress (in the US), or the British Library (in the UK) etc. to claim copyright?
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Project Unica is an initiative by the University of Illinois libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
How long until someone starts making fake rare books to sell to AI companies?
Looking at my modest shelf of weird midcentury travel logs, I just want to be spared from this moralizing. This is just another case of everyone seeing store shelves as an extension of themselves, just like the decline of other physical media. If their knowledge was so precious then why was it not on the used bookstores and libraries to save them, especially when scanning has become so accessible in the last decade? Why did I find some of these books rotting in overpacked shelves and boxes?
Do you think they will make all of that publicly available after it's been scanned?
Being purchased and juiced for model weights is about as noble of an end as any book could hope for.
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
What often gets missed is that they are buy one physical copy and turning it into a digital copy.
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
Aren't AI companies all about the rare book auctions now?
This reads as bad propaganda. Surprised that they left the "translated from Chinese" disclaimer. There is the false implication here that the AI companies are somehow tracking down all copies of a particular book and destroying it. AI has spawned a global freak-out.
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
This is only because this is a legal right granted as part of the purchase of copyrighted material, and because we have tried to stop AI companies from doing this with purely digital copies.
The insanity of attempting to prevent AI learning (which is a direct consequence of the nature of observable information) because of the myth of intellectual property is the main driver of this type of behavior.
I don’t get this latest anti AI talking point. They are digitizing the books preserving them forever.. are you upset that you don’t have access to it? Because you didn’t before either… stop whining and give AA some money.
>It’s outrageous is that it’s legally permissible
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
I hate that to save this article about an AI company acting abhorently I have to 'favourite' it.
Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.
Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
What’s really disgusting is how unnecessary this is.
LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.
And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.
During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions
"Whoever destroys a book destroys a link in the chain of human knowledge"
-- Thos. Jefferson
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
again "destroy" is not accurate. to scan massive amounts of books you have to cut the binding part as feed scanning has been around for ages
there are non destructive scanning options but they are nowhere as fast and prone to errors
Does the demonization of destruction of books mean I shouldn't delete downloaded e-books off my phone's Kindle app? All those precious bits, gone forever...