Normal view

Someone Is Mysteriously Snapping Up Used Books Around the World

5 August 2026 at 20:40

Last week, the internet raged as social-media posts and news articles accused AI companies of “destroying the world’s books”—including “millions of rare” ones—by chopping off their spines to scan them more easily, and discarding them afterward. One article said the news was “sparking concerns that the last remaining copies of out-of-print texts are being destroyed.” The investor and AI skeptic Michael Burry called the practice of sacrificing rare books “evil incarnate.” Even Elon Musk weighed in, posting that he had asked the engineers training xAI’s models to “preserve any rare books in a library and scan them the hard way.”

The destruction of physical books in the process of training AI has been public knowledge for more than a year, but debate over the practice intensified after an article by 404 Media, which described a large number of orders coming to booksellers, suggested that rare books might be included, and named a book-database company called ISBNdb as the possible buyer.

Lost in the outrage, however, are several unanswered questions. The 404 Media story said ISBNdb had advertised bulk buying of books for AI companies, but acknowledged that no evidence directly connects the company to the recent wave of purchases that booksellers have described. And although the story calls out rare-book sellers, it’s unclear how rare most of the books being bought actually are, or whether they’re of great value.

Despite that, the conversation rapidly intensified into a panic. So what is actually going on? And who is actually responsible?

[Read: ]Is This What Comes After AI Slop?

Many of the specifics of this situation are obscured by hard-to-trace purchasing histories and the silence of AI companies and other potential buyers, so let’s lay out what we know. We know that at least one major AI company has used physical books to train their models. Documents unsealed earlier this year in a lawsuit against Anthropic reveal that in the spring of 2024, the AI giant bought millions of new and used books and destroyed them in order to scan the pages to train its AI models. Anthropic and other AI companies are likely keen on books published before the launch of ChatGPT, in 2022, which are all but guaranteed to be free of AI-generated text. Although AI-generated text is readable—if often unpleasant—to humans, it can be devastating to AI models: When trained on their own output, they can degrade significantly in quality in a phenomenon known as “model collapse.”

We also know that something mysterious is afoot. Someone—if not more than one group or company—has been snapping up massive volumes of books from shops around the world. “What’s with the online purchasing bonanza of used books right now?” asked one bookseller on Reddit three months ago. Several people who identified themselves as booksellers responded with their own stories of unusually large orders.

Charlie Becker, the manager of Becker’s Books, in Houston, told me that 95 of the past 100 books ordered through one of his online-selling platforms went to a single buyer. One of the orders was for 70 books. (To avoid breaking the terms of service of his selling platform, Becker would not tell me the name of the buyer.)

But contrary to many news reports, the old books bought from these sellers do not appear to be “rare” in the typical sense of the word. According to Becker and the online discussions I read, the books in these orders are almost entirely nonfiction, were mostly published from the 1970s to the 1990s, and were written in a variety of languages. Becker wrote in a blog post that he has received orders for “stuff like The Insider’s Guide to Metro Denver from 1995 or How to Use Corel WordPerfect 1991.” These belong to a category of books that you could call rare because a small number of copies remain, but they may also be out-of-date, undesirable to readers, and not valuable per se. In some cases, booksellers have noted, the cost of shipping exceeds the books’ resale value.

By most accounts, AI companies want a large volume of text for as little cost as possible. It would be strange if they were intentionally paying premium prices for first editions, unusual printings of cult classics, or yellowed books of esoteric knowledge. Yet it’s possible: The pursuit of training data that their competitors don’t have could lead them to such titles.

We also have a good sense of who’s behind the book buying—and it seems unlikely to be ISBNdb.com, the subject of the 404 Media story. For one thing, no bookseller I could find in online discussions had mentioned an order from ISBNdb. The company maintains a public book database and charges for access to it, but does not appear to buy or sell any books. Things get confusing because, in March, the company added a page to its website that offered AI developers the ability to buy physical books, “up to 1 million titles per order,” with a nondisclosure agreement to preserve anonymity.

When I asked ISBNdb about the service, I received an email from a customer-support representative claiming that “ISBNdb has never purchased, scanned, or destroyed a book—for AI training or anything else. We don’t train AI models, and we never have. The page was up to explore demand for a service we never brought to life.” In response to recent “concern,” the representative said, the company had removed the page advertising that service from its website.

Reddit users did call out a different company for “frequent and sizable online orders”: Zoom Books, based in Canada, claims to use “smart logistics” to “acquire, sort, and resell over 2 million books monthly.” Booksellers in Germany mentioned similar orders from Zoom, as did a New Zealand bookseller who wrote that Zoom had placed multiple orders, including one for 71 titles. Becker would not confirm or deny whether his bulk buyer was Zoom.

On May 26, a book-industry newsletter called Publishers Lunch noted accusations that Zoom was buying books for AI companies. On the following day, the newsletter included a response from Zoom, claiming that it “has no involvement in digitization, scanning, or destruction of books, and we do not sell books for those purposes to our knowledge.” (Zoom did not respond to a request for comment.)

Some booksellers who received orders from Zoom have reported that the shipping addresses on these orders belong to a company called PrepFort. PrepFort specializes in preparing items for sale on Amazon and Walmart, but as a third-party logistics company, it can route books to other destinations as well. Using such a company would be consistent with the AI industry’s practice of acquiring training data through intermediaries, a strategy that has been used for images, paywalled articles, and other media. I emailed PrepFort to see if it could provide any details about these orders. Sufyaan Kalota, a co-founder of the company, responded: “While we would love to assist, we unfortunately cannot share any details regarding our customer agreements due to our confidentiality commitments.”

In June, the NL Times, an Amsterdam-based publication, reported that some Dutch booksellers had received emails from a Singapore-based organization called 2077AI, which claims to be “revolutionizing AI data,” partly by creating training data sets. One bookseller reported that the email from 2077AI was accompanied by a list of 3,000 English-language titles the company wanted to purchase. The examples given were academic books such as Distinct Element Modelling in Geomechanics and Laser Shock Peening of Advanced Ceramics. Such books might legitimately be considered “rare”: they had small print runs, contain esoteric knowledge, but are important within a specific field. (404 Media linked to this story but did not reference these books specifically, or 2077AI.) 2077AI did not immediately respond to a request for comment.

For now, this is what we know: A lot of used books are suddenly being bought up by companies including Zoom Books. But we do not know how many of these orders are coming from AI companies, or whether these books will be destroyed. Based on the practices of Anthropic and the attempted purchases by 2077AI, it seems very possible that the books are destined for AI companies, but we can’t be certain.

Regardless, the idea of an AI behemoth destroying books has sent a bolt through the collective psyche. Book destruction is historically linked to authoritarian regimes and attempts to control public discourse, as well as ancient tragedies—the burning of the Library of Alexandria comes to mind. The encroachment of AI, particularly in the realm of books, is now everyday news: Last month it emerged that a top-selling Kindle book may have been partly written by AI, and grandparents are buying children’s books filled with AI-generated slop.

Even if a literal book burning is not under way, the process of ingesting millions of books, stripping them of authorship, and blending them into a homogenous “intelligence” branded with the name of a chatbot certainly feels destructive. It masks the hard work and collaboration that goes into knowledge creation, undermines incentives for authors to write books, prevents experts from finding one another, and gives AI companies tremendous power over what information people can access. Given this scale of cultural damage, the physical destruction of books seems almost quaint.

© Illustration by Alisa Gao / The Atlantic

Why Would Meta Download So Much Porn?

24 July 2026 at 21:15

Editor’s note: This work is part of AI Watchdog, The Atlantic’s ongoing investigation into the generative-AI industry.


Amid tech companies’ ongoing efforts to vacuum up as much data as possible to train AI models, Meta has taken public posts from its own platforms, scraped massive amounts of content from the rest of the internet, and pirated millions of books to train its AI models. Now the company is being accused of downloading much more illicit material: huge volumes of porn, nonconsensual celebrity nudes, blueprints for 3-D-printable handguns, and millions of passwords acquired by hackers.

Strike 3, the parent company of Vixen Media Group, a prolific producer of adult films, is suing Meta for allegedly downloading 2,973 of its copyrighted videos. But in its efforts to collect evidence about these purported downloads, Strike 3 also captured information relating to many other files that Meta may have acquired over the course of two years: They include images from “Celebgate,” a 2014 hack that resulted in the leak of private photos belonging to Jennifer Lawrence, Kirsten Dunst, and others, plus several collections of deepfake celebrity porn featuring the faces of Natalie Portman, Scarlett Johansson, Elizabeth Olsen, Gal Gadot, and others.

Strike 3’s file-transfer lists include ordinary movies and TV shows, alongside dozens of videos from GirlsDoPorn, which was shut down after six people associated with the site were charged with sex trafficking. Meta allegedly downloaded several individual episodes from the website and two “GirlsDoPorn MegaPack” archives. Strike 3 also presents evidence that Meta made some content publicly available in addition to downloading it.

When I reached out to Meta, a spokesperson told me via email that “these claims are bogus.” Meta also noted in a court filing that Strike 3, which regularly files lawsuits against people who download its work illegally, “has been labeled by some as a ‘copyright troll.’”

This is a complicated case. Piracy can be hard to track with precision, and both parties are clearly conflicted: Strike 3, in seeking damages, wants Meta to look bad, and Meta would prefer not to be associated with this kind of content. Yet if Strike 3’s data are accurate, they would show that participating in mass piracy is a by-product of modern tech development—whether or not any of the material collected is explicitly used to engineer AI models or other programs.

[Read: The hypocrisy at the heart of the AI industry]

Although Meta acknowledged the possibility that employees, contractors, or even visitors to Meta’s offices may have used the company’s networks to download the porn in question, its defense in court was that the material was accessed for “personal consumption” rather than as part of a company project. But neither the sheer volume nor the pattern of downloads looks like personal consumption. The file-transfer logs produced by Strike 3 show sequences of files that are unrelated except for a certain keyword or concept. For example, two files that were transferred consecutively on June 2, 2023—“Bajillion Dollar Properties S01E07” and “First Time Home Buyer Anal Fantasy”—are both loosely themed around real estate, but the similarities likely end there. Judge Eumi Lee, who denied Meta’s motion to dismiss the case, cited other such juxtapositions, such as consecutive downloads of “Teenage Mutant Ninja Turtles (1987-1996)” and a video labeled “Teen Sex Sessions 2 (2012).”

Strike 3 has argued that the content could be used for AI training. A Meta spokesperson told me, “We don’t want this type of content, and we take deliberate steps to avoid training on this kind of material.” In its legal defense, Meta pointed out that its terms of service prohibit users from using its AI products to generate images containing pornography. Even if Meta isn’t building a porn-generating bot, however, tech companies can use “unsafe” content to test guardrails for their systems—effectively telling the software what not to generate. (Lee noted that Meta’s terms of service are irrelevant to the question of whether it wants to acquire porn.)

It’s also possible that Meta is compiling an archive of material for no precise purpose, just as Anthropic acquired millions of books that it says it had no intention of using for AI training but wanted to keep for its “research library.” As AI-training techniques advance, there is no telling what kind of content a company might find useful in the future.

Nineteen of the allegedly downloaded files contain models of functional handguns that can be produced with 3-D printers. A file labeled “5.7 million passwords list” purportedly contains 5,718,107 passwords from accounts that were hacked from 2015 to 2019. The downloads also include password-cracking tools that can be used by hackers to break into people’s accounts. Other files include the movies BlacKkKlansman, The Banshees of Inisherin, and everything made by Studio Ghibli from 1979 to 2020; music by Dua Lipa, Elton John, and Paul McCartney; radio shows and podcasts such as The Howard Stern Show and 1619; and software including Microsoft Windows, Adobe Photoshop, Ableton Live, and NBA 2K23.

The files were downloaded through BitTorrent, a system that allows people to share files among themselves. When you download files via BitTorrent, you typically share those files with other people at the same time; if nobody is sharing the files, then nobody can download them. According to Strike 3’s data, which lawyers from the company told me it regularly gathers in an effort to track the illegal distribution of its intellectual property, devices on Meta’s networks also made some of Strike 3’s videos available to others—a practice known as “seeding”—for as long as three months after fully downloading them.

Meta additionally used BitTorrent to acquire large quantities of copyrighted books around the same time, as shown in testimony from another lawsuit. Some Meta employees were uncomfortable with this activity. According to court documents, Jelmer van der Linde, a former Meta employee who was asked to download books to train the company’s AI models in 2024, told his supervisor in a work chat message that it seemed like a “very very dark grey area. Legally, ethically …” He added that torrenting “copyrighted material is not very legal either.” He was transferred to a different project, and another developer downloaded the books instead. He has since left Meta and now works at the Ellison Institute of Technology in Oxford.

[Read: The unbelievable scale of AI’s pirated-books problem]

To verify Strike 3’s data, I reached out to Tom Chothia, a professor of cybersecurity at the University of Birmingham who has studied BitTorrent. He reviewed a technical description of Strike 3’s data-collection methods and told me that, assuming the records weren’t tampered with, “they have established that people at Meta were uploading” Strike 3’s videos.

One thing that complicates the case is that most of the torrenting did not occur over Meta’s corporate networks. Strike 3’s data show sequences of related file transfers across several networks, including Meta’s. Strike 3 argues that the other networks were used remotely in a coordinated way, perhaps to make the activity harder to trace. Meta has called the patterns a coincidence. But Lee found them compelling, writing that Meta’s coincidence theory “strains belief” and that the activity is likely indicative of “algorithmically coordinated behavior.”

Nearly 300 of the transfers were traced to a private residence in Mountain View, California. Strike 3 claims that this was the home of a Meta contractor’s father, and that the file transfers stopped when the contractor stopped working for Meta. In response, Meta argued that these downloads were “plainly indicative of personal consumption.” According to Strike 3’s logs, the files downloaded to the residence include, in addition to porn, Russian-language books, Chinese movies, and a wide variety of cracked software, among other things.

Unfortunately, in the AI era, illegal mass-downloading by tech companies is not uncommon. Other lawsuits have revealed that OpenAI and Anthropic have also used BitTorrent to download copyrighted books. Whether they have uploaded those books as well, or what other content they may have downloaded through BitTorrent, is not yet clear.

BitTorrent is a strange world in which to find respectable megacorporations. It is dominated by media piracy, and very little legitimate file-sharing occurs. But if AI models are intended to replace human creators, then they have to be trained on as much human-created content as possible. BitTorrent is perhaps the fastest and cheapest way to download high-quality work in large quantities. AI companies have claimed that their models can create original work, but so far, the development of artificial intelligence has involved a lot of theft of other people’s intelligence.

© Illustration by Akshita Chandra / The Atlantic

❌