The rise of artificial intelligence has brought an unexpected casualty: the physical book. To feed massive language models, some AI companies have been purchasing second-hand books in bulk, only to destroy them for digital scanning. This practice, while efficient for data collection, is raising alarms among bibliophiles and archivists who fear irreplaceable works could disappear before they are ever digitized.
Why scanning rare books is urgent
Many rare and out-of-print books exist only in physical form, held in private collections or small libraries. When AI firms buy these volumes, they often cut off the spines and run pages through high-speed scanners, a process that destroys the original. While the digital copy may preserve the text, the physical artifact—its binding, marginalia, and historical context—is lost.
Librarians and preservationists argue that a coordinated effort to scan rare books before they are bought up by AI companies is essential. Public archives and non-profits are stepping in, creating digitization projects that prioritize works most at risk. These initiatives aim to build comprehensive digital libraries that are accessible to researchers and the public alike, ensuring that knowledge is not locked away in proprietary datasets.
The tension between AI development and cultural preservation is not easily resolved. On one hand, vast amounts of text are needed to train sophisticated models. On the other, the destruction of physical books represents a loss of heritage that cannot be undone. The hope is that by raising awareness and accelerating digitization, we can save these treasures before it's too late.
For now, the message from the preservation community is clear: scan first, scan fast. Every book digitized is a safeguard against the irreversible loss of our written legacy.
Comments
No comments yet.