AI companies polluted the internet with slop. Now they are harvesting and destroying millions of rare books
In January, the Washington Post reported:
In early 2024, executives at artificial intelligence start-up Anthropic ramped up an ambitious project they sought to keep quiet. “Project Panama is our effort to destructively scan all the books in the world,” an internal planning document unsealed in legal filings last week said. “We don’t want it to be known that we are working on this.”
Within about a year, according to the filings, the company had spent tens of millions of dollars to acquire and slice the spines off millions of books, before scanning their pages to feed more knowledge into the AI models behind products such as its popular chatbot, Claude.
Details of Project Panama, which have not been previously reported, emerged in more than 4,000 pages of documents in a copyright lawsuit brought by book authors against Anthropic, which has been valued by investors at $183 billion. The company agreed to pay $1.5 billion to settle the case in August, but a district judge’s decision last week to unseal a slew of documents in the case more fully revealed Anthropic’s zealous pursuit of books.
The new documents, along with earlier filings in other copyright cases against AI companies, show the lengths to which tech firms such as Anthropic, Meta, Google and OpenAI went to obtain colossal troves of data with which to “train” their software.
The Anthropic case was part of a wave of lawsuits brought against AI companies by authors, artists, photographers and news outlets. Filings in the cases show top tech firms in a frantic, sometimes clandestine race to acquire the collected works of humanity. [Continue reading…]
Pieter de Vries, an antiquarian bookseller in Haarlem, recently received an email from a person identifying herself as Nataly from Singapore-based company 2077AI, Dutch news site BNR reported.
The woman wrote that the company was undertaking a “a new project focused on collecting books in multiple languages, currently mainly in English”. “We have compiled a very extensive list of editions that we are currently trying to acquire, and we plan to place a fairly large order,” the email said.
The attached list contained 3000 English-language titles organised by ISBN number, ranging from books on fairytales and folklore to technical manuals and science texts.
Media outlets in the Netherlands, Switzerland, Spain, and Germany report booksellers have received nearly identical requests, per NL Times.
Booksellers believe the purchase requests were unrelated to collecting or resale due to the obscure, highly specialised nature of the titles, concluding they were being used to train AI models.
Large language models (LLMs) that power AI tools like ChatGPT until now have mainly been trained on vast datasets that include internet data, licensed books and articles, code repositories and human feedback.
But with much of the content available online now exhausted — and increasingly polluted with poor-quality, AI-generated writing — AI labs are turning to uncorrupted texts published pre-2022.
One company, ISBNdb, which boasts that it has the “world’s largest book database”, now offers bulk book buying for AI labs, “up to one million titles per order”, including of “older, rare and specialist volumes”.
“The world’s best AI training data is sitting on a shelf,” it states on its website. [Continue reading…]