Skip to main content
Back to News
Anthropic has been scanning and discarding millions of books for AI training, a practice that is rai...
Technology
2 min read
US

Anthropic has been scanning and discarding millions of books for AI training, a practice that is rai...

The AMW Read

Updating player map for Anthropic with a data sourcing practice that carries significant IP and heritage implications, though not a groundbreaking shift.
NoveltySignificance
Foundation Models · Player MapData & IP
Anthropic
Anthropic

Foundation Models / LLMs

View Company Profile

Anthropic has been scanning and discarding millions of books for AI training, a practice that is raising concerns about the preservation of rare and antique books. The company reportedly cuts the spines off physical books to feed them through high-speed scanners, a common but destructive method in mass digitization projects. The sheer volume of books processed has alarmed librarians, archivists, and rare book dealers, who worry that irreplaceable works are being lost in the pursuit of training data.

This practice highlights the growing tension between AI companies' insatiable appetite for high-quality text data and the cultural imperative to preserve physical artifacts. While scanning books for AI training is not new—Google Books and the Internet Archive have faced similar criticism—the scale and lack of transparency here are notable. For the AI industry, it underscores the ongoing scramble for training data beyond publicly available web text, pushing companies toward physical archives and proprietary collections. It also raises ethical and legal questions about the destruction of cultural heritage in the service of model development.

For builders and investors, this signals that the data supply chain is becoming a critical bottleneck and a reputational risk. AI labs may need to develop more sustainable data acquisition strategies, such as partnering with libraries for non-destructive digitization or investing in synthetic data generation. Startups and vendors offering data provenance and ethical sourcing solutions could find growing demand. Meanwhile, regulators and cultural institutions may push for stricter oversight of large-scale scanning operations, potentially affecting how Anthropic and others access offline text corpora going forward.

#AItraining #data #Anthropic #books #preservation #ethics

#Anthropic#AI training#data acquisition#preservation#books

How This Connects

Based on Foundation Models · Player Map

  1. 3h agoFour leading Google DeepMind researchers — Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le —...Discovery Loop
  2. 3h agoAnthropic has been scanning and discarding millions of books for AI training, a practice that is rai... · THIS ARTICLE
  3. 19h agoOpenAI has paused parts of the development of its next-generation model, Astra, after internal evalu...OpenAI
  4. 19h agoAnthropic has officially filed for a confidential IPO with the U.S. Securities and Exchange Commissi...Anthropic
  5. 2w agoMoonshot AI's K3 model faces distillation controversy as US officials allege IP theft, rekindling open debate on model training practices.Moonshot AI
  6. 1mo agoIn the Weights launches AI-centric vanity search for personal brand monitoringIn the Weights

Related News

More news from Anthropic

Stay updated with the latest news and announcements from Anthropic.

View all Anthropic news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard