
Anthropic has been scanning and discarding millions of books for AI training, a practice that is rai...
The AMW Read
Updating player map for Anthropic with a data sourcing practice that carries significant IP and heritage implications, though not a groundbreaking shift.
Anthropic has been scanning and discarding millions of books for AI training, a practice that is raising concerns about the preservation of rare and antique books. The company reportedly cuts the spines off physical books to feed them through high-speed scanners, a common but destructive method in mass digitization projects. The sheer volume of books processed has alarmed librarians, archivists, and rare book dealers, who worry that irreplaceable works are being lost in the pursuit of training data.
This practice highlights the growing tension between AI companies' insatiable appetite for high-quality text data and the cultural imperative to preserve physical artifacts. While scanning books for AI training is not new—Google Books and the Internet Archive have faced similar criticism—the scale and lack of transparency here are notable. For the AI industry, it underscores the ongoing scramble for training data beyond publicly available web text, pushing companies toward physical archives and proprietary collections. It also raises ethical and legal questions about the destruction of cultural heritage in the service of model development.
For builders and investors, this signals that the data supply chain is becoming a critical bottleneck and a reputational risk. AI labs may need to develop more sustainable data acquisition strategies, such as partnering with libraries for non-destructive digitization or investing in synthetic data generation. Startups and vendors offering data provenance and ethical sourcing solutions could find growing demand. Meanwhile, regulators and cultural institutions may push for stricter oversight of large-scale scanning operations, potentially affecting how Anthropic and others access offline text corpora going forward.

