I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks
Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.
> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
There's just a few copies in a single library worldwide which is probably a national or a university library
At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
My personal waste book at home is so rare, it's unique. That doesn't mean it needs preservation.
It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
Ai companies don't use ebooks, because they are more expensive than second hand books
They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.
I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks
Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.
To do what?
To not destroy rare books.
They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.
Since when do AI companies care about copyright law? They're destroying them so their competitors can't use them.
Since they had to pay 1.5 billion for it
If only this complaint was being posted by an organization ideologically opposed to copyright itself!
> Instead, they enforce the copyright and force AI companies to shred books they want to ingest. What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
[delayed]
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
[delayed]
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
Since when do AI companies care about the law? Most of their training data is pirated.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
It's giving Vishnu, but the world cannot exist without Shiva.
Aren't AI companies all about the rare book auctions now?
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
No it is not. https://reddit.com/r/Annas_Archive/comments/1vrvt9e/athome_s...