Amazon has been acquiring large quantities of rare and out‑of‑print books, removing their spines and digitizing the pages for use in training its large language models. The activity was traced by 404 Media, which placed a tracker in a rare volume that ended up at Amazon’s VGT3 facility in Las Vegas, marked by a dinosaur‑holding‑a‑book logo.
The company says it obtains these texts through ordinary commercial channels to improve its products. The drive for such material stems from the need to supplement internet‑scraped corpora with content that predates widespread LLM generation, thereby reducing the risk of model collapse that occurs when models are trained heavily on AI‑produced text.
- Acquisition method: spine removal and scanning of rare books
- Facility identifier: VGT3, Las Vegas, dinosaur‑holding‑a‑book emblem
- Stated purpose: improve Amazon products/services via commercially sourced texts
- Underlying rationale: avoid model collapse by using pre‑2022, human‑authored material
Why this matters
Inference: This practice highlights a growing tension between the demand for novel training data and intellectual‑property rights, as sourcing rare physical books may skirt digital copyright limits while raising questions about the legality of reproducing protected works for model training. It also underscores the industry’s reliance on heterogeneous data sources to maintain model robustness as synthetic data accumulates.
