logoalt Hacker News

trollbridgetoday at 1:23 PM7 repliesview on HN

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals).

We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.

We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.

And no, no AI company has ever come to us and asked to run training on all of our scanned copies.


Replies

palmoteatoday at 1:28 PM

> We reprint old books after checking out copyrights

Who is we?

> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely

Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.

show 1 reply
Aurornistoday at 3:27 PM

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies

The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.

There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.

show 1 reply
vander_elsttoday at 2:01 PM

Some source or citation or context is needed here, is this the work of a 2 person no profit or a trillion valued pre IPO company?

show 1 reply
htrptoday at 2:10 PM

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies

Yet

butliketoday at 1:42 PM

Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?

show 3 replies
yorwbatoday at 1:30 PM

Most likely they don't know you exist. If you contacted them first, maybe something could be arranged...

show 1 reply
ck2today at 1:41 PM

or just send the books/scans to a country that doesn't recognize US copyright