Amazon operates a covert book-scanning operation at a Las Vegas facility that strips rare and copyrighted materials to feed AI training datasets, according to tracking data uncovered by researchers. The facility receives bulk orders of books, removes bindings, and scans pages before destroying the physical copies. This process raises serious copyright and intellectual property concerns, particularly for rare publications and out-of-print works that hold historical value.
Researchers planted tracking devices in book shipments and traced deliveries to the Las Vegas location, confirming Amazon's systematic approach to acquiring training data for large language models. The company acquires books through bulk purchasing, often sourcing rare editions and copyrighted works without explicit author or publisher consent. Once scanned, the originals are discarded.
This practice highlights the tension between AI development and copyright protection. Authors and publishers have sued multiple AI companies, including OpenAI and Meta, over unauthorized use of copyrighted material in training datasets. Amazon's operation suggests the tech giant may be bypassing licensing agreements entirely by acquiring physical books and destroying them after extraction.
The discovery comes amid broader scrutiny of AI training practices. Publishers including Penguin Random House have fought legal battles to prevent their catalogs from fueling generative AI models. Some authors report their entire bodies of work appearing in training sets without compensation or attribution.
Amazon has not publicly disclosed this scanning facility or its role in AI training infrastructure. The company's book-buying ecosystem, including Goodreads and direct publishing platforms, gives it unique access to vast literary repositories. This vertical integration allows Amazon to quietly acquire, scan, and destroy materials at scale.
The practice raises questions about whether companies can legally destroy books they purchase for data extraction. It also suggests AI training data sourcing operates in regulatory gray zones where copyright enforcement remains weak. As AI companies race to scale models with diverse training data, book destruction for machine learning represents a troubling new precedent in how legacy intellectual property gets repurposed without consent.