AI companies destroying books face backlash over training data
404 Media reports AI firms are buying and pulping physical books for model training, raising copyright and rare-book concerns.
By Sofia Marchetti · Columnist
· 3 min read
AI companies destroying books has become a flashpoint in the race to build better chatbots, according to reporting by 404 Media. The practice matters beyond publishing: it shows how costly, legally sensitive and operationally messy the AI data race has become for companies trying to improve their models.
404 Media reported that AI companies are using intermediaries to buy physical books in bulk while keeping the buyers anonymous. The books are then taken apart, scanned into digital training datasets and discarded, according to the report.
Bulk-sourcing companies are advertising the ability to find hundreds of thousands of titles for AI clients while promising confidentiality, 404 Media reported. That secrecy reflects the tension around training data, as AI developers face lawsuits and public scrutiny over how they obtain material used to build models.
Why are AI companies destroying books?
AI models need large amounts of text to learn patterns in language. Books published before the broad rise of generative AI in 2023 are especially valuable to developers because they are more likely to be written by humans rather than generated by machines, according to critics cited by 404 Media.
Generative AI means software that can produce new text, images, code or other output after being trained on large datasets. Critics sometimes call low-quality AI-generated material “AI slop,” and they argue that developers are trying to capture human-written books before newer text collections become more polluted by machine-made writing.
The demand is already changing parts of the used-book business. One unnamed bookseller told 404 Media that weekly sales rose from about 20 books to several hundred after AI-linked buyers entered the market.
The bookseller said the buying has helped financially and cleared inventory that might otherwise sit unsold, especially overseas inventory and foreign-language books. Still, he told 404 Media he dislikes the end use and is concerned that uncommon books are being pulped after scanning.
What courts have said so far
The practice resembles Anthropic’s reported “Project Panama,” which digitized millions of books through destructive scanning, according to International Business Times coverage cited in the reporting.
In the copyright case Bartz v. Anthropic PBC, a federal judge in San Francisco ruled last summer that scanning legally purchased physical books into digital copies could qualify as transformative fair use, even when the originals were destroyed. Fair use is a U.S. copyright doctrine that can allow certain uses of protected works without permission, and “transformative” use generally means the new use serves a different purpose or character.
Federal judges later issued similar fair-use rulings in separate copyright cases involving OpenAI and Meta, according to Decrypt. Those decisions did not end the broader fight over AI training data.
In a separate case in the same federal district, a judge this week approved a $1.5 billion copyright settlement requiring Anthropic to pay thousands of authors about $3,000 per book after the company used pirated copies of their works to train Claude, according to ABC News.
Industry backlash is growing
The criticism is not only coming from authors and booksellers. Elon Musk wrote on X that he had asked the SpaceXAI team to preserve rare books in a library and scan them without cutting off the spine.
For investors watching the AI sector, the takeaway is that data is becoming a core input with legal, reputational and supply-chain risks attached. Better models may need better training material, but the fight over how that material is collected is far from settled.
This story draws on original reporting from Decrypt.