A federal judge in California has allowed some of the most damaging claims against Meta Platforms to proceed in a copyright lawsuit that alleges the tech giant didn’t just download pirated books to train its AI models — it seeded them back out to other users on BitTorrent. The ruling, issued by U.S. District Judge Vince Chhabria, strips away several of Meta’s defenses and sets the stage for what could become the most consequential copyright case in the AI training debate.
The implications are staggering.
If Meta’s engineers did in fact participate in BitTorrent file-sharing networks to acquire copyrighted training data for its LLaMA large language models, and if those engineers left their BitTorrent clients running in default mode — which automatically redistributes downloaded files to other users — then Meta wasn’t merely a passive consumer of pirated material. It was an active distributor. That distinction matters enormously under copyright law, where willful distribution of protected works carries far steeper penalties than mere unauthorized reproduction.
The case, brought by authors including Richard Kadrey and Sarah Silverman, has been winding through the Northern District of California since 2023. But as TorrentFreak reported, the latest ruling from Judge Chhabria represents a significant escalation. The court found that plaintiffs had adequately alleged that Meta engaged in the unauthorized distribution of their works — not just unauthorized copying — through the BitTorrent protocol’s built-in seeding function.
Here’s the technical background that makes this so damaging for Meta. BitTorrent is a peer-to-peer file-sharing protocol. When a user downloads a file, the default behavior of virtually every BitTorrent client is to simultaneously upload — or “seed” — pieces of that file to other users who are also downloading it. This is what makes the protocol efficient: every downloader becomes an uploader. To avoid seeding, a user must deliberately disable the feature or close the client immediately after the download completes.
The plaintiffs allege Meta’s employees used BitTorrent to download massive datasets of copyrighted books, including collections sourced from shadow libraries like LibGen and Z-Library. And they allege Meta did nothing to prevent the default seeding behavior, meaning Meta’s servers were actively distributing those pirated books to strangers on the internet while the downloads were in progress.
Meta’s lawyers tried to get these distribution claims thrown out. Their arguments, as characterized by the court, were unpersuasive. According to TorrentFreak, Meta argued that the plaintiffs hadn’t sufficiently shown that anyone actually received the files Meta allegedly seeded. Judge Chhabria rejected this, noting that the very nature of BitTorrent means that seeding inherently involves distribution to other users — that’s how the protocol works. The judge found that alleging Meta used BitTorrent with its default settings was enough to plausibly claim distribution occurred.
Meta also argued that even if seeding happened, it was merely incidental to the downloading process and shouldn’t count as a separate act of infringement. The court wasn’t buying it. Distribution is distribution.
The ruling is particularly notable for what it says about the credibility gap Meta faces. Internal communications and evidence already surfaced in the litigation suggest that Meta employees were well aware they were obtaining copyrighted material through unauthorized channels. One key piece of evidence involves an internal Meta communication where an employee reportedly discussed the use of LibGen datasets — one of the largest repositories of pirated academic papers and books on the internet. The fact that employees apparently knew the provenance of these datasets undercuts any future fair use defense, because it speaks to the willfulness of the conduct.
And willfulness is where the financial exposure gets serious. Under the Copyright Act, statutory damages for willful infringement can reach $150,000 per work infringed. With potentially thousands of copyrighted books at issue, the theoretical damages could reach into the billions. That’s before considering the distribution claims, which could multiply the exposure further.
This case doesn’t exist in isolation. Meta is one of several AI companies facing copyright litigation from authors, publishers, visual artists, and news organizations. OpenAI is defending against suits from The New York Times and several authors. Stability AI faces claims from Getty Images and visual artists. But the BitTorrent seeding allegation is unique to Meta and introduces a dimension of liability that other AI copyright cases lack. Downloading copyrighted material without authorization is bad enough. Redistributing it to the world is categorically worse.
The broader AI industry is watching closely. The question of whether training AI models on copyrighted data constitutes fair use remains unresolved. Multiple cases are working through the courts, and no appellate court has ruled definitively on the issue. But Meta’s situation is more precarious than most, because the seeding allegations could be adjudicated separately from the fair use question entirely. Even if a court were to eventually find that using copyrighted data for AI training qualifies as fair use — a big if — that ruling wouldn’t necessarily protect Meta from liability for distributing those same works via BitTorrent.
Think of it this way: a researcher might have a plausible fair use argument for photocopying a book for personal study. But that argument evaporates if the researcher then hands out copies on a street corner. BitTorrent seeding is the digital equivalent of handing out copies on a street corner.
Judge Chhabria did dismiss some claims in the case. He tossed allegations related to certain statutory provisions that the plaintiffs hadn’t adequately supported. But the core copyright infringement claims — both for unauthorized reproduction and unauthorized distribution — survived. So did claims under the Digital Millennium Copyright Act related to the removal of copyright management information. The plaintiffs allege that Meta stripped metadata and copyright notices from the books before feeding them into its training pipeline, which is a separate violation under the DMCA.
Meta has consistently maintained that its use of publicly available data to train AI models is protected by fair use. The company has pointed to the transformative nature of AI training, arguing that the models don’t reproduce copyrighted works but instead learn patterns and relationships from them. This is the same argument advanced by virtually every AI company facing copyright claims, and it has some legal support in prior cases involving search engines and text mining. But those precedents involved fundamentally different technologies and use cases, and courts have signaled they won’t simply rubber-stamp fair use claims in the AI context.
The seeding issue makes Meta’s position uniquely difficult to defend. Fair use is a four-factor balancing test under Section 107 of the Copyright Act, and one of those factors is the effect of the use on the market for the original work. If Meta was redistributing pirated copies of books to other BitTorrent users, it was directly contributing to the availability of those pirated copies — which directly undermines the market for legitimate sales. That’s about as bad as it gets under the fourth factor.
There’s also the question of what this means for Meta’s relationships with publishers and content creators going forward. Several major publishers have already signed licensing deals with AI companies — Wiley with OpenAI, the Associated Press with OpenAI, and others. These deals are premised on the idea that AI companies need permission and should pay for access to copyrighted content. If courts find that Meta helped itself to pirated books and then seeded them back out, it poisons the well for any future licensing negotiations. Why would a publisher negotiate in good faith with a company that allegedly pirated its catalog?
The timing of the ruling also matters. Congress is actively considering legislation to address AI and copyright, with multiple bills introduced in both the House and Senate. Lawmakers have held hearings featuring testimony from authors, musicians, publishers, and AI executives. A finding that one of the world’s largest technology companies used BitTorrent to pirate and distribute copyrighted books would significantly strengthen the hand of those pushing for stricter regulation of AI training data practices.
Meta declined to comment on the specifics of the ruling when reached by multiple outlets. The company has previously stated that it believes its AI training practices are lawful and that it respects intellectual property rights.
For the authors who brought the case, the ruling is a vindication of their theory that Meta’s conduct went far beyond the usual AI training copyright dispute. Richard Kadrey, Sarah Silverman, and their co-plaintiffs have argued from the beginning that Meta’s use of pirated datasets was qualitatively different from, say, scraping text from the open web. The BitTorrent seeding allegations sharpen that distinction to a razor’s edge.
What happens next will likely involve discovery — the legal process through which both sides exchange evidence. The plaintiffs will almost certainly seek Meta’s internal records related to the downloading and use of BitTorrent, including server logs, employee communications, and technical documentation about how the training datasets were assembled. If those records confirm that seeding occurred and that Meta employees were aware of it, settling the case before trial could become Meta’s most rational option. But any settlement large enough to satisfy the plaintiffs would set a precedent that other rights holders would immediately seek to replicate.
So Meta faces a classic litigation dilemma: fight and risk a devastating trial verdict that establishes binding precedent, or settle and invite a flood of similar claims. Neither option is attractive.
The case number is Kadrey v. Meta Platforms, Inc., No. 3:23-cv-03417, in the U.S. District Court for the Northern District of California. It’s one to watch — not just for what it says about Meta, but for what it reveals about the true costs of building AI systems on foundations of pirated content.


WebProNews is an iEntry Publication