Something strange is happening in the world of secondhand books.
Booksellers in the UK, Europe and the United States have reported unusually large orders for old books. Sometimes hundreds or even thousands of apparently unrelated titles are purchased at once.
At first, nobody was quite sure why.
Now there is growing evidence that at least some of these books are being bought to help develop artificial intelligence.
But there is a controversial twist.
Some of the books are being cut apart, scanned and then destroyed.
Why does AI need books?
AI chatbots learn about language by analysing enormous amounts of text.
Books can be particularly valuable because they contain carefully written, edited and organised information. Novels also contain long stories, conversations and different styles of writing.
Older books may offer another advantage.
Books printed before the recent explosion in generative AI are much more likely to contain text written entirely by humans.
The internet now contains a growing amount of AI generated material. Developers therefore have good reasons to value reliable collections of human writing when creating future AI systems.
Millions of printed books represent an enormous collection of human knowledge and language.
What is destructive scanning?
Turning a warehouse full of books into digital information presents a practical problem.
Opening every book and carefully photographing each page would take a long time.
Destructive scanning makes the process much faster.
The binding is removed and the pages are separated. They can then be passed rapidly through industrial scanners.
Once this happens, the original book cannot simply be returned to a bookshelf. The damaged pages can then be discarded or recycled.
It is efficient, but it has understandably upset people who value books as physical objects.
Project Panama
Much of what we know about the practice emerged from a copyright case involving Anthropic, the company behind the Claude AI assistant.
Court documents revealed an internal project called Project Panama.
One Anthropic document described the project as an effort to “destructively scan all the books in the world”.
Anthropic purchased large numbers of physical books and arranged for them to be scanned. According to the court, the bindings were removed, the pages were cut and scanned, and the original paper copies were discarded.
The digital versions were added to a central library that Anthropic could use when developing its AI systems.
The phrase about scanning all the world’s books described the enormous ambition of the project. It does not mean Anthropic literally acquired every book in existence.
It is not just Anthropic under scrutiny
The story became even more interesting in August 2026.
Journalists at 404 Media worked with a bookseller who had received an unusual order for around 1,000 books. A tracking device was hidden inside one of the books.
The shipment eventually arrived at an Amazon facility in Las Vegas.
Workers told the publication that books at the facility were having their bindings removed so their pages could be scanned.
Amazon told 404 Media that it purchases books through commercial channels to help develop and improve products and services used by its customers.
Meanwhile, booksellers in several countries have reported unusually large orders containing seemingly random combinations of titles.
However, we should be careful about assuming that every mysterious bulk order is connected to AI. In many cases, booksellers simply do not know who the final customer is.
But isn’t destroying books legal?
This is where things become complicated.
In the United States, buying a physical book generally gives its owner considerable freedom over what happens to that particular copy.
A 2025 US court ruling involving Anthropic also made an important distinction.
Judge William Alsup ruled that using copyrighted books to train Anthropic’s AI models was a fair use under US copyright law.
He also ruled that converting purchased printed books into digital copies for Anthropic’s internal library was fair use in the circumstances of the case. The company destroyed the printed copy after making its digital replacement.
But there was another part of the story.
Anthropic had also downloaded millions of books from pirate websites. The judge did not give Anthropic the same protection for building this permanent library from pirated copies.
Copyright law also varies between countries.
In the UK, the rules surrounding copying copyrighted material for AI training are different from those in the United States.
This means there is no simple worldwide rule saying AI companies can scan any book they want.
What happens when a rare book disappears?
There is another issue beyond copyright.
Most people probably would not worry too much if one copy of a modern book with millions of copies in circulation was recycled.
Rare books are different.
Some books have survived for hundreds of years. Others were produced in tiny numbers and may contain information that is difficult to find anywhere else.
Anthropic says its data acquisition programmes do not buy and destroy rare or antiquarian books.
However, the recent investigation into Amazon involved a bookseller dealing in rare books. This has increased concerns about what could happen when companies purchase books in large quantities.
A computer system might see a book as several hundred thousand useful words.
A historian, collector or reader might see the same object as an important piece of our shared past.
Both things can be true.
A bigger lesson about artificial intelligence
There is an interesting irony in this story.
Some of the world’s most advanced AI systems are hungry for something humans have been producing for centuries.
Good writing.
Books contain carefully constructed arguments, knowledge, imagination, humour, stories and ideas. That makes them enormously valuable when developing machines that work with human language.
AI can bring huge benefits to education, science and society. But developing it responsibly means thinking about where its knowledge comes from and who created that knowledge.
It also means considering what should be preserved.
The challenge is not simply building more powerful AI.
It is finding ways to build it while respecting authors, copyright, culture and the human knowledge that made these remarkable technologies possible in the first place.








