I'm much less upset about the destruction of rare (old, out of print, niche) books by megacorps than I am about the fact that those books can't be digitized by libraries for public access or benefit because the former is legal and the latter is not.
Amazon can buy a dusty ass book from the 1970s, scan, OCR, and index it.
My only option is to buy the stinky old book, if I can find it, and then barely be able to read it because I don't have access to digitization equipment or software and my library finds the book too niche to stock
If you have actually worked with rare books (I own a bunch) you KNOW they smell awful and the fonts are tiny. The library checkout cards are cute and nostalgic but I'd far rather be able to a) magnify the text on a giant screen b) be able to search it c) use text to speech.
Digitization is amazing.
@ehashman The problem is not computers; the problem is that we have now invented the thing that comes after computers, and it is much, *much* worse than computers
@ehashman or sometimes they *have* been digitized but still can't be shared.
@ehashman Increasing scarcity is good for business. Decreasing scarcity is bad for business.
@WestLawns "why not just carefully photograph 200+ pages per book" they say to the disabled woman who can barely move her arms. Good one. Never thought of this. You got rare and brilliant ideas going in these here replies
@ehashman don't library require special license to stock book while AI scalper use retail version without any special license?
@ehashman
Internet Archive is still in this legal limbo.
The entire capitalist system is crashing by its own weight.
Although I generally agree with you, I'm not 'much less upset', merely 'less upset'. Because, as a book lover, the process of destroying a book in order to digitize it bothers me on a visceral level.
And even more so with rare books.
@ehashman agreed, instead it's solely used to feed in information for LLMs, there was no harm in listing those books, it's honestly frustrating where we're heading, my only hope at this point is our generations knowledge, hopefully we can take better action
@ehashman Consider that it's even worse than you think...
A rare book is more than just the letters on a page. There's information about the binding technology used, the ink, the paper processes, the fonts and formatting, printer artefacts.
Who knows what insights about history and past technology we may have been able to extract (non destructively) from the physical book, instead of feeding it into a mulcher.
@ehashman Ich muss in letzter Zeit oft an Aaron Swartz denken.
@ehashman @internetarchive @internetarchiveeurope I do wonder if you could provide a public training set for LLMs under the same fair use clause. Make it easy to index and ingest by offering it in all your usual formats.
@ehashman Same. The Internet Archive has been under legal fire for years because they do what libraries do (and what they are registered as) and lend books they actually own, and publishers decided that's not how ownership works and they should purchase a separate digital license and pay a fee again every time someone reads said book. But scanning them, making commercial derivatives of their contents and destroying them? completely fine.
@ehashman
If the #NinthGate actually existed, #Amazon would find, scan, and destroy all copies.
@ehashman The data is scanned, making the book available in digital form is simply a matter of legislation
@ehashman
The real issue is that by scanning the books you can feed the whole books to an llm and that llm generates a low quality version of that....
@wooramel are you aware that most of these books are mass-printed 20th century books that are out of print? They're not hand-written manuscripts, nearly all the coverage of this practice has talked about books with an ISBN
Much prefer physical copies of old books.
@davefischer that's nice, I have a neurovisual disability and I don't because I can't actually read them.
@ehashman right? They are gonna be destroyed by time anyways, but feeding them to an llm to be lost forever but part of the collective is dystopian behavior.
@ehashman Sure. I get that.
Someone commented that the destruction is getting value out of books that have been collecting dust on shelves.
I get that.
And in 200 years how many of those books will still be around? What scanning technology will be available in 200 years to examine the books? To get an insight into... atmospheric conditions? Pollen types? Tree DNA mutations?
The unique thing about books is the technology to use them is built into the device. Books are usable for as long as they exist.
How much other information storage technology becomes impossible to use because the readers no longer exist? Laser-disk, anyone?
The real crime is that technology exists to *non-destructively* scan the words in the book, but, you know... nah.
@wooramel the quality of non-destructive scans is generally worse and slower, though. I personally don't have any qualms with destructive scanning so long as the information is preserved for public use. The problem is, it's not, and can't be due to the way intellectual property law forbids this for works in copyright.
@ehashman I mean, how bad does a scan have to be for the OCR to fail? The problem for AI companies is that it's slower and more expensive.
Scanning the text only gets the text. There is more information in a physical object than just the text. Once that's gone, it's gone forever. It prevents the future 200 year tech from discovering more information about when the book was produced.
My flatmate likes classic cars. At one point they were mass produced. If the only way to understand the tires on classic cars was to shred them he'd be livid, incandescent with rage. We all should be.
Once the last Morris 1000 is gone, it's gone forever. When the last copy of first edition Watership Down is gone, it's gone forever. All the information about the technology and incidentals of the time goes with it.
And, as I say, there are perfectly fine non-destructive methods to get *exactly* the same result the AI monsters are after.
@ehashman Someone in these scanning orgs needs to leak the whole trove to Anna's Archive. Since the LLM providers used the shadow libraries to train, they have a duty to give back.
Does anyone else think this book scanning operation is largely legal cover for using shadow libraries? They can now say "oh we scanned the book" if accused.
@ehashman To add to this , I can't even buy the books because shipping on a physical item has only been getting more and more costly, snd nine times out of ten the only available source is not in the same country as me.
The shipping cost often significantly exceeds the cost of the book itself. So mostly I just give up. (Usually they're sewing or pattern making books, but there's also been a few on leatherworking, woodworking, and other such niche and obscure crafts, so they're often not going to get uploaded anywhere.)
Multiply by entire craft communities, and there's the situation of the only access being via the handful of people who either snagged a copy of a particularly helpful book when it was new and more affordable (but they usually don't have a complete scan, or any scans at all, usually due to thinking the FBI will come smash their door down if they even share verbatim text of a page) or have the disposable income to spend hundreds of dollars on a single book. And when those people either pass away or the worst happens to their collection... the place it doesn't end up is digitized and publicly available.
@ehashman As I said on Lemmy, this is the second coming of Nazi book burnings.
@ehashman yes! This is enclosure of the intellectual commons
@dartigen @ehashman And who gets to store that for 200 years?
What about the detail scanning technology in 200 years would have caught that we're missing now?
The book that is being destructively scanned now will be gone forever. The data will only last, say, 40 - 50 years before the database applications no longer work or are supported or exist or the storage tech will be obsoleted or there will be a mass data attack which will wipe out a bazzillion books or *whatever* you can't predict now.
The 'original scan' is the physical book. You're right in that having that is crucial.
@wooramel @ehashman Yeah, but nobody with the means to correctly store a physical book for 200 years has the interest in doing so broadly enough to save everything that might be important to someone someday.
Even storing a small collection of books with a view towards preservation is *expensive*. People with that kind of money are rare in craft communities, and are usually more focused on the craft. (And IME tend to not really understand how the books can be critical to people trying to get into that craft, but that's an argument for a different thread.)
Scans aren't perfect, sure - but they're better than a pile of rotting paper that's no longer readable.
And that is the fate that befalls a lot of the book collectors in these communities - their houses burn down or get flooded out, they don't think about issues like mold or insect damage... and even if they do, I can't begin to tell you how many of them likely don't have detailed instructions in their will about what needs to happen to their books, so they'll probably end up in thrift shops anyway. (Or even have wills. Or families that will actually understand the importance and follow the instructions properly. And even then...)
But if someone had scanned that book, the information is still there. And if they shared that scan, that's two people who have that information (or 10, or 50...) and can store it at a much lower cost to themselves. (Of course these AI publishers won't be sharing the scans at all, so they're also not helping.) Hell of a lot easier to share a scan too - I don't have to pay international shipping on a book, or get on a plane, and everyone else who wants access to it can still access it. And if even 10 people have downloaded that scan, then if my hard drive dies and loses everything, that's 9 other people with a copy.
It'd be nice if libraries had the funding to be able to store everything, even niche self-published craft books, but they don't, and that doesn't seem likely to change. And leaving that task up to individuals and subject-focused communities isn't exactly working out well.
@ehashman I'm still spinning from Copyright Law just *disappearing* overnight.
@claralistensprechen3rd @dartigen USPS media mail is still a thing but it's for domestic shipping only, and there was some talk of abolishing it recently (which would have been very bad for interlibrary loans)
@claralistensprechen3rd Project Gutenberg only covers books in the public domain. The rare books that these megacorps are purchasing are under copyright.
@ehashman
I also prefer reading text on a screen, but destroying books is a no-no. These are separate issues.
@hajovonta books do not last forever and are regularly destroyed, including by your favourite library and booksellers. The destruction of a single book does not destroy the information contained within, nor other printings of the same book.
Both are illegal, technically.
Sorry, let me clarify.
1. If the book has extant copyright, it is illegal to convert it to a different format, whether Amazon does it or a library does.
2. If the book is old and out of copyright, as it seems the current books that Amazon is doing this to are, then a library could very well buy it and scan it and do the same. Anyone could. There would be no prohibition on converting it in that case.
@jmcrookston 1. was explicitly ruled as fair use in the case against Anthropic. https://deadline.com/wp-content/uploads/2025/06/anthropic.pdf
> the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central
library with more convenient space-saving and searchable digital copies for its central
library β without adding new copies, creating new works, or redistributing existing copies.
@jmcrookston Libraries can't do this because the purpose of scanning would be to distribute new copies, which is not permitted.
I would assume essentially all books being discussed here are under copyright because if they weren't, there wouldn't be any need for the destructive scanning process.
I'm not aware of anything legally prohibiting a library from scanning an out of copyright book and making it available.
If it's within copyright, then Amazon cannot do it either. Of course, Amazon may not care, but that's a separate issue.
I'll have to review the Anthropic decision you have posted, which I just noticed now.
It's a first instance summary judgment ruling that held yes, it was fair use (exception to copyright) for Anthropic to convert books to digital format. There is at least one other similar case, see https://www.lexology.com/library/detail.aspx?g=674f228b-9c25-48c0-8c24-51e6df425674.
I would wait to see what happens at the appellate level because this is inconsistent with copyright law generally. However, it is certainly what this judge held in this case. A library could argue it is entitled to do the same, presumably.
@jmcrookston The case was settled so I don't think there are any appeals planned: https://www.insidetechlaw.com/blog/2025/09/bartz-v-anthropic-settlement-reached-after-landmark-summary-judgment-and-class-certification
That's a shame. This case took a very expansive view of the fair use exception.
Friend of mine built and sold DIW book digitisers a while ago. Dunno if they still do, but the concept was: Canon camera with free firmware, place to place the book, you press a pedal which pushes a piece of glass onto the book to hold the open pages in place and triggers the camera, you release the pedal, flip the page, rinse, repeat.
Allowed you to scan a while book into a PDF in an hour at most.
Absent that, a standard scanner should also work, though be slower?
- replies
- 0
- announces
- 0
- likes
- 0
@ehashman that is not quite correct. Especially books of scientific interest get digitized as fast as libraries can do it. The problem is more that digitalization of old books a manual process if you don't want to destroy them and therefore a very resource intensive endeavour. E.g. at the university of Heidelberg ~10 people work full-time on digitalization. If copyright allows it, which is usually the case for everything of interest, they get published for public access.
These books are not digitised and then add available electronically to the public.
They are digitised, fed into the LLM training corpus, and destroyed.
You will not be able to get either the physical nor the digital version.