AI Companies AI Companies

AI Companies Are Buying Rare Books Just to Destroy Them | Here’s What They’re Really After

Artificial intelligence companies need data on an almost unimaginable scale.

Websites, academic papers, photographs, code, videos and books can all become valuable sources for training increasingly sophisticated AI systems. But there is one enormous collection of human knowledge that cannot simply be downloaded from the internet: books that were never properly digitized.

That problem has created an unusual new market.

AI-related companies and data providers are reportedly acquiring physical books in bulk, removing their bindings and scanning their pages so the text can be converted into machine-readable datasets. The process is known as destructive scanning, and the name is surprisingly literal.

The book may have to be physically cut apart.

Once that happens, putting it back on a bookshelf in its original condition is generally impossible.

The practice has attracted particular attention because some books being sought are old, obscure or difficult to replace. Yet from an AI developer’s perspective, those forgotten volumes contain something extremely valuable: language and information that may never have appeared on the public web.

So why would anyone destroy a perfectly usable book to teach a computer?

AI Has Already Consumed an Enormous Amount of the Internet

Large language models learn statistical patterns from enormous quantities of text.

The internet initially provided an obvious source.

Web pages contain news, reference material, discussions, technical documentation, fiction, reviews and countless other forms of human writing.

But web data has limitations.

It contains duplication.

It contains spam.

It contains machine-generated material.

And increasingly, it contains text produced by AI itself.

Researchers have raised concerns that repeatedly training models on synthetic material could degrade performance under some circumstances, particularly when synthetic data begins replacing the diversity found in human-created datasets. A 2024 study published in Nature demonstrated a phenomenon researchers described as model collapse when models were recursively trained on model-generated data.

Books offer something different.

They contain long-form, edited human writing created before generative AI existed.

That makes libraries look increasingly like enormous stores of relatively clean human-generated data.

Millions of Books Still Aren’t Properly Available Online

It is easy to assume that almost every published book has already been digitized.

That is not true.

Major digitization projects have scanned millions of volumes, but countless books remain difficult to obtain electronically.

Some had tiny print runs.

Others are decades or centuries old.

Some were published by regional presses that disappeared.

Technical manuals, local histories, specialized reference works and forgotten fiction can exist almost exclusively as physical copies.

Organizations such as the nonprofit Internet Archive have spent years preserving and digitizing physical books, illustrating both the scale and difficulty of converting printed collections into searchable digital information.

For an AI company, a physical-only book represents previously inaccessible training material.

And previously inaccessible data can be unusually valuable.

Why Can’t They Just Photograph Every Page?

They can.

But it is slow.

Preservation-focused digitization is designed to protect the original object. A book may be placed carefully in a specialized scanner while pages are photographed individually without damaging the binding.

That makes sense when the physical book itself needs to survive.

It makes much less economic sense when someone wants to process hundreds of thousands of inexpensive books as quickly as possible.

Destructive scanning changes the equation.

The spine or binding can be removed so individual pages become loose sheets. Those sheets can then pass through high-speed document scanners capable of processing pages much faster than someone manually turning them.

The images are subsequently processed using optical character recognition, or OCR, which converts photographs of printed words into machine-readable text.

Once the text has been extracted and checked, the physical pages may have little remaining value to the company that purchased them.

The book has effectively been transformed from an object into data.

Why Are Old Books Particularly Attractive?

Consider what exists inside a book published in 1962.

Its vocabulary comes from humans.

Its sentence structures come from humans.

Its illustrations were created before generative image systems.

Its facts reflect the knowledge available at that time.

And, crucially, none of its content was influenced by ChatGPT or other modern generative systems.

Older books therefore offer AI developers a kind of pre-AI snapshot of human culture.

That can be useful when developers want training material uncontaminated by synthetic text.

Books also cover subjects the modern internet barely discusses.

An obscure 1950s engineering manual may explain equipment no longer manufactured.

A regional history may document people and places poorly represented online.

An old scientific text can show how terminology evolved.

A forgotten novel contributes literary structures, dialogue and vocabulary different from contemporary web writing.

Each individual book may appear unimportant.

Across millions of volumes, however, those differences create enormous linguistic and informational diversity.

Rare Doesn’t Necessarily Mean Valuable

The phrase “rare books” can create the wrong mental picture.

AI companies are generally not expected to be slicing apart priceless first editions of Shakespeare or Gutenberg Bibles.

In the bookselling world, rarity and monetary value are different concepts.

A book can be difficult to find without being particularly valuable.

A technical manual printed 60 years ago may have only a handful of surviving copies available for sale, yet collectors may have almost no interest in it.

That could make the book inexpensive enough to purchase for scanning while still making its contents rare digitally.

This distinction matters.

For AI training, an obscure $12 book whose text does not exist online can potentially be more interesting than a famous $5,000 first edition whose text has already been digitized hundreds of times.

The AI company wants informational scarcity.

The collector wants physical scarcity.

Those are not always the same thing.

The Real Commodity Is Human-Written Text

The strange economics make more sense when the book is viewed as a data container.

Suppose a company pays $5 for a used book containing 300 pages.

That is potentially tens of thousands of words of professionally edited human writing.

If the company can automatically scan, OCR, clean and index the material, the acquisition cost per word becomes extremely small.

Now repeat that process across 100,000 books.

The company has created a massive corpus containing billions of words from sources that may be poorly represented in existing online datasets.

The physical books were simply the delivery mechanism.

This helps explain why AI has changed the economic value of previously overlooked information.

A warehouse full of unwanted books once looked like a storage problem.

It can now look like raw material.

Copyright Doesn’t Disappear When Someone Buys the Book

Buying a physical book does not automatically transfer its copyright.

Someone purchasing a novel owns that physical copy.

He does not normally acquire the right to reproduce and commercially distribute the author’s work.

That distinction has become central to the legal battle over AI training.

Authors and publishers have sued AI companies over the use of copyrighted books, arguing that developers copied protected works without authorization.

AI companies have generally argued that training can qualify as fair use because models analyze works to learn statistical relationships rather than simply distributing identical copies.

U.S. courts are now beginning to address those arguments directly. The U.S. Copyright Office has also examined generative-AI training as part of its broader AI initiative. Its materials and reports can be explored through the U.S. Copyright Office AI initiative.

Importantly, buying a legitimate physical copy can change some legal questions surrounding acquisition, but it does not magically eliminate copyright law.

The legality of copying and using the contents for AI training remains a separate issue.

Anthropic Put Destructive Book Scanning Into the Spotlight

One of the most important legal cases exposing this practice involved Anthropic.

In litigation brought by authors, court records described Anthropic purchasing millions of print books and having their bindings removed so the pages could be scanned into digital form for its central research library.

In a major June 2025 ruling, U.S. District Judge William Alsup concluded that Anthropic’s use of legally purchased books for training was fair use in the circumstances before the court. However, the judge treated Anthropic’s separate acquisition and retention of pirated books differently. The distinction between legitimately purchased physical books and unlawfully obtained copies became central to the ruling.

That decision attracted enormous attention because it demonstrated why purchasing and destructively scanning physical books could have legal significance beyond simple convenience.

A legitimately acquired book provides a very different provenance story from a file downloaded from a piracy site.

Why Not Simply License Everything From Publishers?

Licensing sounds like the obvious solution.

In some cases, it may be.

Publishers control enormous digital catalogs and could theoretically provide high-quality electronic files without destroying anything.

But licensing every relevant book would be complicated.

Rights can be fragmented.

Publishers disappear.

Contracts expire.

Authors regain rights.

Different rights may exist in different countries.

Some old books have unclear ownership histories.

Others remain protected by copyright even though no obvious rights holder can easily be found.

Then there is cost.

AI companies want datasets containing enormous numbers of works. Negotiating individual agreements across that scale could become extremely expensive and slow.

That does not resolve the legal or ethical debate.

It explains why developers have strong incentives to find alternative routes to large collections.

Destroying Books Feels Different Because Books Are Cultural Objects

There is another reason the practice provokes unusually strong reactions.

People do not treat books like ordinary storage devices.

A damaged hard drive is electronic waste.

A destroyed book can feel like lost history.

Books carry annotations.

They contain inscriptions.

They have printing characteristics that reveal when and how they were produced.

They can contain marginal notes from previous owners.

Even the binding and paper can tell historians something about publishing practices.

Organizations devoted to preservation therefore treat digitization differently from companies interested primarily in extracting text.

The Library of Congress maintains extensive preservation programs precisely because the physical artifact can have cultural value beyond the words printed on its pages.

A perfect text transcription preserves the sentences.

It does not necessarily preserve the object.

The Biggest Concern Is Irreplaceable Material

Destroying one copy of a mass-market paperback printed five million times is unlikely to erase cultural history.

Destroying one of only several surviving copies of an obscure publication is different.

That is where critics of indiscriminate destructive scanning have a legitimate concern.

Large-scale acquisition systems may evaluate books primarily according to whether their contents are useful or unavailable digitally.

Preservation requires another question:

How many physical copies still exist?

Unfortunately, that can be difficult to determine.

Library catalogs document institutional holdings, but privately owned copies may be invisible.

A book that appears common could be disappearing.

A book that appears rare could have hundreds of copies sitting in private collections.

Responsible digitization therefore requires more than purchasing whatever used books happen to be available.

It requires understanding what should be preserved before deciding what can safely be sacrificed.

There Is an Irony in Destroying Books to Preserve Their Knowledge

The practice creates a strange contradiction.

A company destroys the physical book.

But by scanning it, the company may create one of the most durable digital copies of its contents ever produced.

An obscure volume sitting in someone’s attic can be destroyed by moisture, insects, fire or simple neglect.

Once digitized, its text can theoretically be copied across multiple storage systems indefinitely.

So is destructive scanning preservation or destruction?

It can be both.

The words survive.

The object does not.

That distinction has existed in library science for decades, but AI has dramatically increased the scale at which it matters.

AI Has Given Forgotten Books a Completely New Kind of Value

The broader story is not really about companies suddenly developing an appetite for destroying literature.

It is about the growing value of authentic human-generated information.

The early internet provided AI developers with enormous quantities of it almost for free.

As models improve and AI-generated material spreads across the web, high-quality human-created datasets become increasingly valuable.

Books represent one of the largest remaining reservoirs.

They contain carefully edited prose.

They cover centuries of knowledge.

Many predate the internet.

Almost all older books predate generative AI.

And enormous numbers have never been fully digitized.

That turns a dusty used bookstore, library discard pile or warehouse of forgotten publications into something the AI economy increasingly understands as data infrastructure.

The Book Isn’t What the AI Company Wants

That is ultimately the key to understanding the entire practice.

An AI company buying thousands of old books is not necessarily building a library in the traditional sense.

It wants what is trapped inside them.

Language.

Facts.

Writing styles.

Historical perspectives.

Technical explanations.

Stories.

Human reasoning recorded before machines began generating large portions of the internet.

Cutting the binding allows that information to move from paper into a machine-readable dataset much faster.

From an engineering perspective, it is efficient.

From a preservation perspective, it can be unsettling.

And from a copyright perspective, it sits inside one of the most consequential legal battles of the AI era.

The strange result is that artificial intelligence has made some of the oldest information technology humans possess the printed book valuable all over again.

Not necessarily because anyone wants to read the physical copy.

But because machines increasingly need what humans wrote inside it.

Leave a Reply

Your email address will not be published. Required fields are marked *