A book is usually thought of as a physical object: something to read, preserve, lend, sell or archive. But in the race to build increasingly powerful artificial intelligence systems, a book can also be something else: a dataset.
A recent investigation into bulk purchases of used and rare books in the United States has brought this transformation into sharp focus. A bookseller placed an Apple AirTag inside a book from a large anonymous order. The tracker eventually led investigators to an Amazon facility in Las Vegas where workers reportedly process books through a destructive scanning operation: cataloguing them, removing their bindings, scanning the pages and disposing of the physical remains.
Amazon has acknowledged purchasing books through commercial channels, saying that the purchases help develop and improve its products and services. It has not, however, publicly confirmed that the particular operation identified in the investigation is being used to train artificial intelligence models.
That distinction matters.
The story is not simply about Amazon, booksellers or the future of publishing. It raises a much bigger question about data governance in the age of AI: what happens when the world’s physical knowledge is converted into machine-readable data, and who gets to decide how that knowledge is acquired, transformed and controlled?
Table of Contents
- The Strange New Economics of Books
- Why Are AI Companies Interested in Old Books?
- From Library to Dataset
- The Copyright Question Is Only Part of the Story
- Africa Has an Even Bigger Question to Ask
- Digitisation Should Not Mean Destruction
- The African AI Governance Conversation Must Include Cultural Data
- The Risk of a New Form of Data Extraction
- What Should African Institutions Do?
- The Real Lesson From Las Vegas
The Strange New Economics of Books
According to the investigation, booksellers had been observing unusual purchasing patterns for more than a year.
Anonymous buyers were purchasing hundreds or even thousands of books at a time, often targeting obscure academic works, older non-fiction titles and out-of-print material. These were not necessarily valuable first editions. Instead, buyers appeared interested in the information contained within the books.
The AirTag investigation provided a physical trail.
The tracked book travelled from California through distribution facilities before ultimately reaching Amazon’s LAS8 facility in North Las Vegas. There, investigators connected the shipment to a unit known as VGT3, where workers reportedly scan large quantities of books. Employees described a process involving barcode or ISBN scanning, removal of bindings and high-speed scanning of individual pages.
In other words, the book’s physical existence is temporary.
Its value, from the perspective of the process, is in the information that can be extracted from it.
Once the text has been digitised, the original object may have little further value to the system.
That is a profound change in how society treats knowledge.
Why Are AI Companies Interested in Old Books?
The answer lies partly in the growing scarcity of high-quality training data.
Large language models require enormous quantities of text. For years, the internet provided an apparently endless supply of publicly accessible material. Books, websites, academic papers, news articles, forums and other digital content could all contribute to training datasets.
But the internet is changing.
Generative AI now produces an enormous volume of text, creating a difficult problem for future AI training: distinguishing human-generated material from machine-generated material.
If models repeatedly train on content produced by previous models, errors and distortions can be reinforced across generations. Researchers have described this broader phenomenon as a risk of model collapse.
This makes older human-authored material increasingly valuable.
Books published before the widespread emergence of generative AI provide a relatively stable pool of human-created text. And unlike much web content, books can contain specialised knowledge that was never digitised or easily accessible online.
That helps explain the interest in obscure academic and out-of-print works.
The AI industry may not need the book.
It needs the words inside it.
From Library to Dataset
This is where the story becomes particularly important for data governance.
Digitisation is not inherently harmful. Libraries have been scanning books for decades. Digital archives can dramatically expand access to knowledge, protect fragile materials and make historical collections searchable.
The issue is control.
Who is digitising the material? For what purpose? Under what legal authority? Where will the resulting dataset be stored? Who can access it? Can the dataset be sold or licensed? Can it be used to train commercial AI systems? And what happens to the original physical material after digitisation?
These questions are familiar to anyone working in data protection and information governance. But the book-scanning controversy demonstrates that they extend well beyond personal data.
Data governance is increasingly about knowledge governance.
The Copyright Question Is Only Part of the Story
The legal debate surrounding AI training has largely focused on copyright.
In the United States, a 2025 ruling involving Anthropic found that using lawfully purchased books to create a training dataset could qualify as fair use because of the transformative nature of the use. The legal position is not universal, however, and other jurisdictions take different approaches.
Copyright therefore matters enormously.
But copyright does not answer every governance question.
Imagine an AI company legally purchases the only surviving physical copy of a rare African publication. It scans the book, extracts the text and destroys the original.
The copyright question might have one answer. The cultural heritage question could have another. The archival question could have another. And the sovereignty question could be different again.
A legal right to acquire an object does not necessarily resolve whether society should lose that object forever.
Africa Has an Even Bigger Question to Ask
For Africa, the story should prompt a much broader conversation.
Across the continent, enormous quantities of knowledge remain trapped in physical form.
They exist in:
- university libraries;
- government archives;
- newspapers;
- research institutions;
- museums;
- community collections;
- historical records;
- indigenous knowledge repositories;
- old legal and policy documents;
- religious archives; and
- out-of-print African publications.
Much of this material has never been comprehensively digitised.
That creates both an opportunity and a vulnerability.
AI companies seeking high-quality human-authored material could increasingly find value in these collections. The question is whether African institutions will digitise, govern and preserve their knowledge on their own terms, or whether valuable physical collections will become another source of data extracted into systems controlled elsewhere.
This is where digital sovereignty becomes more than a slogan.
If African knowledge is converted into datasets that are hosted, processed and monetised outside the continent, African societies could lose visibility and control over how their intellectual heritage is used.
The physical book may disappear while its digital contents become part of a proprietary model that Africans cannot inspect, access or govern.
Digitisation Should Not Mean Destruction
There is an important distinction between digitising knowledge and destroying its physical source.
For fragile materials, digitisation can be a preservation strategy. But destructive scanning introduces a different risk.
Once a unique physical object has been cut apart and discarded, digitisation cannot fully reverse the loss.
A digital copy captures text. It may not capture everything that makes an object historically meaningful.
A physical book can contain handwritten annotations, ownership marks, marginalia, printing characteristics, paper quality, illustrations, bindings, stamps and evidence of how the object travelled through communities.
Those features can themselves be historical data.
A dataset may preserve the words while erasing the context.
That is why archives and libraries cannot be treated simply as warehouses of training material.
They are custodians of collective memory.
The African AI Governance Conversation Must Include Cultural Data
Africa’s emerging AI governance frameworks increasingly focus on issues such as algorithmic accountability, privacy, cybersecurity, transparency, bias and human rights.
Those remain essential.
But the Las Vegas story suggests another category deserves attention: cultural and intellectual data governance.
African governments, universities, libraries and cultural institutions should be asking:
- What knowledge is being digitised?
- Who owns the resulting digital representation?
- Who controls access to it?
- Can it be used to train commercial AI systems?
- Are communities whose knowledge is represented being recognised or compensated?
- Are culturally significant materials being preserved physically even after digitisation?
- Where are the resulting datasets stored and processed?
These questions become particularly important for indigenous and community knowledge, historical records and works that may fall into complicated areas of copyright or collective cultural ownership.
The Risk of a New Form of Data Extraction
Africa has experienced many forms of resource extraction.
The concern now is that the next frontier could be intellectual.
Instead of extracting minerals, companies extract information. Instead of shipping physical resources, they acquire books, records, images, audio, cultural material and datasets.
The value is then realised elsewhere through technologies, platforms and intellectual property.
This does not mean every foreign acquisition or digitisation project is exploitative. Nor does the Las Vegas investigation prove that African books are currently being systematically acquired and destroyed for AI training.
But it does expose a model worth watching.
The underlying pattern is simple:
Acquire knowledge → digitise it → convert it into machine-readable data → build proprietary systems → monetise the resulting intelligence.
Without appropriate governance, the original knowledge holders may have little visibility into the final use.
What Should African Institutions Do?
The immediate response should not be to block digitisation.
Instead, African institutions should build stronger governance around it.
1. Create national and institutional digitisation policies
Digitisation projects should have clear rules covering ownership, licensing, access, preservation, commercial use and downstream AI applications.
2. Establish provenance requirements
Where datasets are created from books, archives or cultural collections, institutions should maintain records showing where the material originated and how it was processed.
3. Protect culturally significant collections
Libraries, museums and archives should identify materials whose destruction would constitute an irreversible cultural loss.
Digitisation should complement preservation, not automatically replace it.
4. Negotiate AI licensing deliberately
Where valuable collections are licensed for AI development, institutions should consider whether licences should permit model training, commercialisation, redistribution or further data extraction.
5. Build African data infrastructure
The more African knowledge is hosted and governed through African institutions, the greater the ability to determine how that knowledge is accessed and reused.
6. Treat knowledge as an economic asset
African intellectual and cultural collections should not be viewed merely as historical repositories. They are also potential sources of research, education, innovation and economic value.
Governance should therefore protect them accordingly.
The Real Lesson From Las Vegas
The most unsettling part of the story is not the paper shredder.
It is the possibility that we are entering an era where physical knowledge is increasingly valuable only because it can be converted into digital intelligence.
A rare book can survive for a century. A training dataset can be copied in seconds. A physical archive can belong to a community. A model trained on its contents may belong to a corporation.
That asymmetry deserves serious attention.
The AI industry’s hunger for high-quality data is unlikely to disappear. If anything, competition for reliable human-generated knowledge will intensify.
The question for Africa is whether its institutions will simply become suppliers in that emerging data economy, or whether they will develop the governance, infrastructure and bargaining power necessary to determine how African knowledge enters the AI age.
The story unfolding in Las Vegas is therefore not just about what happens to old books.
It is a warning about what happens when knowledge becomes data, and data becomes power.
For Data Governance Africa, that is the question worth following.