AI’s New Hunger for Books: Are Rare Libraries Becoming the Next Casualty of the Artificial Intelligence Race?

Artificial Intelligence has transformed the way we search, write, learn, and communicate. But behind every intelligent chatbot lies an enormous amount of training data. For years, AI companies relied on internet articles, websites, public datasets, and licensed material. Now, according to multiple recent reports, the industry’s search for fresh, high-quality data has moved from the internet to physical bookshelves.

The latest controversy involves a practice known as “destructive scanning”—where printed books are purchased, their bindings are cut off, every page is scanned, and the physical copies are discarded.

While the practice is legal in some jurisdictions under certain circumstances, it has triggered an international debate about copyright, cultural preservation, and the future of printed knowledge.


What Is Destructive Scanning?

Destructive scanning is a digitization technique used to scan books quickly.

Instead of carefully photographing each page while preserving the book, companies:

  • Buy physical books in large quantities.
  • Slice off the book’s spine using industrial cutting machines.
  • Feed the loose pages into high-speed scanners.
  • Convert every page into digital text using OCR (Optical Character Recognition).
  • Dispose of or recycle the remaining paper.

This process permanently destroys the physical copy but produces high-quality digital text suitable for AI training. Reports and court filings involving Anthropic described such efforts under an internal initiative called “Project Panama.”


Why Are AI Companies Buying Physical Books?

The answer lies in the changing quality of internet data.

Large Language Models (LLMs) such as ChatGPT, Claude, Gemini, and others have already consumed enormous amounts of publicly available online information. Researchers now face two major challenges:

1. Internet data is becoming saturated

Much of the useful public web has already been incorporated into AI training datasets.

2. AI-generated content is contaminating the internet

As AI-generated articles, blogs, and websites become more common, training future AI systems on this material can reduce output quality—a phenomenon sometimes called “model collapse.” Older printed books, especially those published before widespread AI-generated content, offer cleaner human-written text.


Why Are Rare and Out-of-Print Books in Demand?

Booksellers across the United States and Europe report receiving unusually large bulk orders for:

  • Rare books
  • Academic publications
  • Technical manuals
  • Specialist reference works
  • Out-of-print editions

Some sellers who previously sold only a few dozen books each week reported suddenly selling hundreds, particularly obscure titles. Dutch booksellers also described requests for thousands of English-language books from companies involved in large-scale collection projects, although the ultimate buyers are not always publicly identified.


Why Are These Books Valuable to AI?

Printed books provide several advantages over web pages:

  • Professionally edited content
  • Accurate grammar and structure
  • Deep subject expertise
  • Long-form reasoning
  • Minimal AI-generated contamination

Books often contain decades of scholarship and knowledge that may never have been digitized, making them attractive as training material for advanced AI models.


The Ethical Questions

The issue extends beyond technology into ethics and cultural preservation.

Loss of Cultural Heritage

A printed book is more than text. Marginal notes, inscriptions, original bindings, and historical context can all have research value. If rare copies are destroyed, those physical features are lost forever, even if the text survives digitally. Antiquarian booksellers have warned that uncommon editions could gradually disappear.

Copyright Concerns

Authors and publishers argue that books should not become AI training material without permission or compensation. AI companies have already faced lawsuits over the use of copyrighted works. Some media organizations and publishers are now calling for licensing agreements rather than unapproved use of creative content.

Fair Use vs. Fair Compensation

In the United States, some court decisions have recognized certain scanning practices as fair use in specific contexts. However, that does not end the broader debate about whether creators should be paid when their works contribute to commercial AI systems. Legal standards also vary across countries.


Not Everyone Supports Destructive Scanning

Criticism has come from authors, publishers, booksellers, and preservationists.

Many argue that:

  • Rare books should be preserved.
  • Libraries should not lose unique editions.
  • Authors deserve licensing revenue.
  • AI companies should prioritize lawful partnerships instead of destroying scarce printed works.

At the same time, some booksellers acknowledge that bulk purchases help clear old inventory and provide financial benefits, illustrating the complex incentives involved.


Could This Affect Libraries?

There is no evidence that AI companies are systematically destroying public library collections. However, preservation experts worry that if rare books become valuable primarily as AI training material, demand for scarce editions could reduce their availability in the second-hand market and private collections. These concerns remain part of an ongoing public debate rather than an established outcome.


What Happens Next?

The controversy is likely to influence future AI regulation worldwide.

Governments and industry groups are increasingly discussing:

  • Copyright licensing for AI training
  • Protection of authors’ rights
  • Preservation of rare books
  • Transparency about training datasets
  • Ethical standards for acquiring source material

Publishers and media organizations are also negotiating licensing agreements that would allow AI companies to access high-quality content while compensating creators.


Final Thoughts

Artificial intelligence depends on knowledge, and books remain among humanity’s richest sources of it. Yet the reported practice of destructive scanning highlights a difficult question: How should society balance rapid AI innovation with the preservation of cultural heritage and the rights of authors?

Digitizing books can help preserve and expand access to knowledge, but when the process involves destroying rare physical copies or using copyrighted works without clear permission, it raises concerns that go beyond technology. The challenge for policymakers, publishers, AI developers, and readers will be finding a path that supports innovation while protecting the literary and historical record that made that innovation possible.

More Reading

Post navigation

Leave a Comment

Leave a Reply

Your email address will not be published. Required fields are marked *