In short
Amazon is reportedly buying rare books, cutting off their bindings and scanning them to create AI training data. The practice highlights the growing scramble for clean, human-written text and raises fresh concerns about preservation and ethics.
- Amazon is reportedly sourcing rare books for AI training data.
- The books are said to be cut apart and scanned at a Las Vegas facility.
- Rare, pre-2022 texts are valuable because they are less likely to contain AI-generated content.
- The practice intensifies concerns about preservation, copyright and model collapse.
Amazon is buying rare books, removing their bindings and scanning them to feed artificial intelligence systems, a move that underscores how aggressively major tech companies are pursuing new sources of training data. The practice matters because it could put fragile, out-of-print works at risk while helping Amazon build better AI tools from text that is difficult to obtain elsewhere.
The issue came into view after 404 Media reported that it tracked a rare book with a device and found it had arrived at an Amazon facility in Las Vegas known as VGT3. Amazon said it purchases books through commercial channels to improve the products and services customers use, but the revelation has prompted fresh scrutiny over how AI companies are sourcing material for large language models.
What Amazon is doing with rare books
Amazon is apparently acquiring physical copies of rare books, dismantling them, and scanning their pages for machine-learning use. That workflow is notable not just because it involves books, but because these are the kinds of texts least likely to exist in a usable digital form.
According to the report, the books are being routed through an Amazon facility in Las Vegas identified as VGT3. The warehouse reportedly uses a dinosaur holding a book as its symbol, a detail that makes the operation easy to identify once the facility is known, but does little to clarify how many books are being processed or exactly which AI systems they will train.
Amazon has not publicly laid out a detailed account of the program. In its statement to 404 Media, the company said only that it buys books through commercial channels for the purpose of improving customer-facing products and services. That phrasing leaves open a broad range of uses, including search, recommendation systems and generative AI products.
Why rare books have become valuable AI fuel
Rare books matter to AI developers because the internet is no longer an inexhaustible source of fresh, high-quality text. Much of the web has already been scraped, filtered and reused by multiple AI labs, and the pool of clean, human-written content keeps shrinking relative to the scale of data these systems require.
Books that are out of print, hard to buy, or absent from major digital collections are especially attractive because they contain long-form writing that has not been repeatedly recycled online. For model builders, that makes them useful for improving vocabulary, style, factual breadth and domain coverage.
There is also a timing advantage. Texts published before 2022 carry a particular value because they were almost certainly written before today’s generation of large language models could have influenced them. That makes them less likely to contain the kind of machine-generated material that can contaminate training sets.
How does AI training data get contaminated?
AI training data gets contaminated when models are trained on text produced by other models, creating a feedback loop that can degrade output quality over time. Researchers often describe this risk as model collapse, a problem in which systems become less accurate, less varied and more repetitive as synthetic text makes up a larger share of the training mix.
For companies racing to improve generative AI, that risk creates pressure to find genuinely human-authored material. Physical books, especially rare ones that have not been widely digitized or copied, offer a comparatively fresh source of text.
- Original human writing is more useful than AI-generated text.
- Out-of-print works are harder to find in existing online corpora.
- Older books are less likely to include model-generated contamination.
- Long-form prose helps train systems on richer language patterns.
Why the story matters beyond Amazon
This is not just a story about one company and a few rare books. It reflects a broader shift in the AI industry: as the supply of easy digital text runs thin, firms are expanding their search for training data into less obvious and more sensitive places.
That shift has already created disputes over scraping websites, licensing publisher archives and using copyrighted material without permission. If books in private or commercial circulation are being physically altered to create training corpora, the debate broadens from digital rights to preservation, ownership and cultural stewardship.
Rare books are not just information containers. They are artifacts, often with historical, aesthetic and monetary value. Cutting off their spines to scan them is a practical method for digitization, but it also transforms the book from collectible object into raw material.
For librarians, collectors and publishers, the central question is whether the transformation is justified by the supposed benefits of AI improvement. For Amazon, the calculus appears to be that better text access can produce better products, even if the books themselves are physically sacrificed in the process.
How Amazon’s explanation compares with the wider AI industry
Amazon is not alone in the hunt for large-scale training data. Across the AI sector, companies have spent years building models on massive text collections assembled from books, web pages, academic papers, code repositories and licensed databases. The challenge is that the best-known sources are increasingly overused or legally contested.
Anthropic, one of the leading AI developers, has already been criticized in relation to books used for training. In that case, allegations centered on the use of pirated works, highlighting how costly and legally complicated it can be to assemble the kind of corpora large language models demand.
Amazon’s approach, based on purchasing physical books through commercial channels, appears different from piracy allegations in at least one respect: it involves buying copies rather than unlawfully copying file collections. But it still raises unresolved issues about consent, destruction and the fate of culturally significant objects once they are repurposed for machine learning.
Amazon said it buys books through commercial channels to improve the products and services customers use, but it did not provide a public explanation for why the volumes must be dismantled or how the scans are used.
What is VGT3 and why does it matter?
VGT3 is the Amazon facility in Las Vegas where the tracked rare book ultimately arrived, according to 404 Media. The site matters because it provides a concrete endpoint for a process that would otherwise be difficult to verify: book acquisition, physical alteration and scanning for model training.
While Amazon has not described the facility in detail, the report’s tracking device suggests a logistics chain that can be monitored and mapped. That kind of evidence is important in a field where companies often disclose only broad categories of use while the operational details remain hidden inside warehouses and data centers.
The Las Vegas location also reinforces how AI development depends not only on software and chips, but on industrial-scale physical infrastructure. Books move from sellers to facilities, pages become images, images become text, and text becomes part of the data pipeline feeding AI systems.
| Key fact | Details | Why it matters |
|---|---|---|
| Company | Amazon | One of the world’s largest technology and retail companies |
| Activity | Buying rare books, cutting off bindings, scanning pages | Creates training data for AI models |
| Facility | VGT3 in Las Vegas | Reported location where a tracked book was found |
| Data value | Out-of-print and pre-2022 texts | Hard to source elsewhere and less likely to be AI-generated |
| Core risk | Model collapse | Training on synthetic text can degrade model quality |
How model collapse changes the race for training data
Model collapse is one of the most important technical reasons companies are searching for new text sources. As synthetic content proliferates across the web, future model training sets risk becoming polluted with machine-written language that may mimic human writing without adding new information.
That can create a feedback loop. One generation of models produces text, the next generation trains on it, and the resulting outputs can become increasingly generic, self-referential or error-prone. In that sense, old books have become more valuable not because they are nostalgic, but because they are clean.
This makes rare books appealing in exactly the way many preservationists would find troubling: they are valuable at scale not as objects to be read, but as reservoirs of text to be extracted.
Why older books may be especially useful
Older books offer a kind of temporal authenticity. They reflect human writing before the rise of mainstream generative AI, making them less likely to include recycled model output. They also represent subjects, styles and voices that may be underrepresented in modern web corpora.
- They may preserve rare language patterns.
- They can include specialized or historical knowledge.
- They often exist outside modern copyright licensing datasets.
- They are harder for competitors to duplicate at scale.
What are the legal and ethical questions?
The legal question is not simply whether Amazon can buy a book. The harder issue is what rights, if any, are implicated when a company acquires a physical copy only to destroy it for data extraction. Owning an object does not automatically resolve broader concerns about reproduction, fair use, licensing or cultural preservation.
The ethical issue is equally complicated. Libraries and collectors preserve books because they are meant to endure, circulate and be studied. Dismantling them for AI training treats the volume as a disposable input rather than a durable artifact.
That tension has become a defining feature of the AI era. Companies argue that access to large and diverse datasets is necessary to build useful tools. Critics counter that the pursuit of scale often overwhelms consent, attribution and preservation.
In Amazon’s case, the company is particularly significant because it sits at the intersection of retail, cloud computing, publishing infrastructure and consumer technology. A data sourcing practice that might seem niche at a startup becomes more consequential when carried out by a global platform.
How the book-scanning process fits Amazon’s broader AI ambitions
Amazon has been investing across AI for years, from cloud services to enterprise tools and consumer-facing features. Training better models requires more than chips and cloud storage; it requires text of the kind that can sharpen those systems’ understanding of language, context and knowledge.
Scanned rare books could support a range of downstream uses, from better search relevance to improved product recommendations and generative features. Amazon has not said which models benefit from the scans, but the logic of the pipeline is familiar across the industry: more diverse and cleaner training data usually means better-performing systems.
The company’s public response remains deliberately broad. By emphasizing that it buys books through commercial channels to improve products and services, Amazon frames the process as a normal part of product development rather than a special AI initiative. But the reported destruction of rare volumes suggests a more complex reality underneath that corporate description.
What happens next?
The immediate next step is likely further scrutiny from journalists, researchers and book preservation advocates. Reports that use tracking devices and facility identification can make hidden workflows visible, which in turn can drive demands for clearer disclosure and more precise policy boundaries.
There may also be renewed debate over whether AI companies should be required to disclose more about the origin of their training data. That discussion has already surfaced around copyrighted books and web scraping, and Amazon’s reported practice could intensify it by adding a preservation angle.
For now, the clearest takeaway is that the AI industry’s thirst for text is not easing. As high-quality online content becomes saturated and synthetic data grows more common, old books may become one of the last rich reservoirs of human writing available to large-scale model builders.
That reality makes Amazon’s reported strategy both practical and unsettling: practical because it solves a data problem, unsettling because it does so by transforming rare physical books into disposable machine-learning input.
Timeline of the reported events
The reporting reveals a straightforward sequence: a rare book was tracked, it reached an Amazon facility, and the broader reason appears to be AI training. The details below summarize the reported progression.
| Stage | Reported development | Significance |
|---|---|---|
| Book acquisition | Amazon buys rare books through commercial channels | Suggests a deliberate sourcing strategy rather than incidental ownership |
| Tracking | 404 Media places a tracking device in a rare book | Allows the destination to be independently verified |
| Facility arrival | The book ends up at VGT3 in Las Vegas | Connects the books to a physical Amazon operation |
| Processing | Books are cut apart and scanned | Converts physical works into digital training data |
| AI use | Scans are used to improve AI models | Links the process to Amazon’s broader AI efforts |
Why this story resonates now
This report arrives at a moment when public trust in AI data practices is already strained. Readers, authors, publishers and researchers have spent years learning that the unseen costs of generative AI can include copyright disputes, environmental strain, labor concerns and the industrialization of data extraction.
Amazon’s reported use of rare books adds a striking visual and cultural dimension to that debate. It is one thing to scrape a website; it is another to imagine a rare volume physically stripped apart in a warehouse so its pages can be ingested by a machine-learning pipeline.
The symbolism is potent because Amazon began as a bookseller. The company’s origin story was built on making books easier to find and buy. This reported practice suggests a very different relationship to books now: not as products to cherish or resell, but as raw material for the next generation of AI.
That shift may prove controversial precisely because it captures the broader transformation underway in tech. The value of text has not disappeared; it has been reclassified. In the AI economy, a rare book may be worth less on a shelf than it is as a source of tokens.
Frequently asked questions
Why is Amazon buying rare books for AI training data?
Amazon is reportedly buying rare books because large language models need huge amounts of high-quality text, and out-of-print works are harder to find online. Older, human-written books are also less likely to contain AI-generated material that could weaken model quality.
What is model collapse in AI?
Model collapse is a decline in AI output quality that can happen when systems train on too much AI-generated text. As synthetic material gets recycled through new models, the results can become flatter, more repetitive and less reliable.
Did Amazon say why the books are being destroyed?
Amazon said it purchases books through commercial channels to improve the products and services customers use, but it did not publicly explain why the volumes need to be dismantled or how the scans are specifically used in AI systems.
Where are the books reportedly being processed?
The books are reportedly being processed at an Amazon facility in Las Vegas known as VGT3. A tracking device placed in one rare book helped connect the physical copy to that site.









