August 31, 2026
Why AI Companies Are Buying Antique Books and What They Are Really After
I’ve always believed that the value of a book is not limited to the information printed on its pages. A book can preserve a way of…

By Alamgir Rajab
4 min read
I've always believed that the value of a book is not limited to the information printed on its pages. A book can preserve a way of thinking, a historical perspective, a vocabulary, or a cultural detail that may disappear from everyday life. That is why I find the recent interest of AI companies in old and secondhand books particularly fascinating.
In recent weeks, booksellers in the United Kingdom, Ireland, Europe, and elsewhere have reported unusual bulk purchases of obscure and out of print books. Some orders contain thousands of titles with little connection between them. The books range from old agricultural manuals and regional histories to racing biographies and specialist nonfiction. The buyers are sometimes difficult to identify, and some sellers suspect that the books are being acquired for artificial intelligence training.
At first, buying antique books for AI may sound strange. Why would a technology company need a physical copy of a book when millions of texts already exist online?
The answer is data.
AI companies are increasingly searching for something more valuable than age or rarity. They are looking for high quality human created information that has not already been absorbed into existing datasets or contaminated by the enormous volume of AI generated material now appearing online.
The Real Product Is the Information
When I look at this trend, I don't see technology companies suddenly becoming collectors of antique literature. I see them becoming collectors of data.
AI models learn from enormous quantities of text. But quantity alone is no longer enough. Researchers have repeatedly emphasized that the quality of training data directly affects the quality of AI systems. As researchers Lora Aroyo, Matthew Lease, Praveen Paritosh, and Mike Schaekermann wrote, "Training data defines what we want our models to learn."
That makes older books surprisingly valuable.
A book published decades ago was created before generative AI existed. It therefore provides human written material that is less likely to contain AI generated language. It can also contain specialized knowledge that may be poorly represented on today's internet.
This matters because researchers have found evidence that repeatedly training AI systems on machine generated content can reduce performance and linguistic diversity. One study found that training on synthetic material can produce significant degradation compared with training on real human generated data.
In other words, some of the oldest information may become useful precisely because it predates the AI content explosion.
Why Physical Books Matter
The interesting part is that companies do not necessarily need antique books because they want to preserve them.
They may want to digitize them.
Recent reporting has described large scale purchases of used books that appear to be intended for scanning. In the Anthropic case, court records described a process in which legally acquired books were stripped from their bindings, scanned, and the physical copies discarded.
The scale of the broader book data opportunity is enormous. Harvard Library researchers recently released a public domain dataset containing 983,004 volumes and approximately 242 billion tokens extracted from books in its collection. The larger Google Books scanning effort covered more than one million volumes.
This changes the way we should think about an old book. Its physical cover may have little value to an AI company. Its contents can be far more valuable.
Antique Does Not Always Mean Rare
There is another important distinction that I think gets lost in the headlines.
The recent buying activity is not necessarily focused on priceless first editions. Reports have described purchases of ordinary secondhand books, including obscure nonfiction published in the late twentieth century. The Atlantic reported that some mysterious orders involved books from the 1970s through the 1990s rather than genuinely rare collectibles.
That tells me the objective is not traditional book collecting.
The objective is coverage.
An AI system benefits from seeing different writing styles, subjects, eras, languages, professions, and perspectives. A forgotten book about local agriculture may contain terminology that a modern website never uses. A regional history may preserve names and events that are difficult to find elsewhere. A technical manual may contain practical knowledge that has never been rewritten for the modern web.
This is where old books become a kind of information reservoir.
The Copyright Question Has Become More Important
Of course, collecting books for AI training raises a difficult question: who has the right to use the information inside them?
The legal situation is complicated and varies by country. In the United States, books published before 1929 are generally in the public domain, although copyright status can require careful verification.
The 2025 Anthropic ruling also demonstrated that the legal debate is not simply about whether AI training can ever qualify as fair use. The court distinguished between legally acquiring books and obtaining copyrighted works through piracy.
That distinction became even more significant in 2026 when a federal judge approved a $1.5 billion settlement involving claims that Anthropic had used pirated books in training its AI systems.
So the future of AI training may increasingly involve not just better models, but better provenance.
What I Think AI Companies Are Really Buying
From my perspective, the most valuable thing inside these books is not nostalgia. It is human knowledge with provenance.
AI companies are operating in an environment where freely available internet data is becoming less reliable as a training resource. The web is increasingly filled with automatically generated articles, summaries, images, and other synthetic material. That creates an unusual situation: AI is helping produce the very data environment that future AI systems may struggle to learn from.
As Judge William Alsup described AI training in the Anthropic case, it could be "quintessentially transformative," while still leaving serious questions around how the underlying books were obtained.
That distinction is important.
The race for AI intelligence is increasingly becoming a race for trustworthy information. The companies that can obtain diverse, human created, legally sourced, well documented data may have an advantage over companies that simply collect more data.
In my book Optimized to Win Creating Content That Ranks, I wrote, "The most successful content is not created for algorithms alone; it is created for people first and optimized second." That principle feels surprisingly relevant here. Even in an age of massive AI models, the original human source remains important.
I see the antique book trend as part of a much larger transformation in the value of information. Books that once sat unnoticed on secondhand shelves may now represent datasets, historical records, specialized knowledge, and linguistic diversity. The real question is no longer why an AI company wants an old book. The more important question is whether society can build systems that allow this knowledge to be used responsibly while preserving access, rewarding creators, and protecting cultural memory.
To sum up, as the founder of a digital marketing agency, I have spent years watching information become the foundation of digital communication. What AI is showing us now is that information does not lose its value simply because it is old. Sometimes, its age is exactly what makes it valuable.
Contact:
President & CEO, Grands Digital
Email: alamgir0500@gmail.com