We Need to Talk About Who Owns the AI Training Data
Big AI labs scraped the internet, got sued, and mostly won or settled their way clear. If you tried the same thing with their content, you would not be so lucky. That asymmetry is the whole story.

The short answer
AI scraping is fair use only sometimes, and mostly for the companies with the lawyers to prove it in court. U.S. courts ruled in 2025 that training a generative model on lawfully acquired copyrighted works can be fair use, but training on pirated copies is not, and the outcome depends heavily on how the data was obtained and whether it competes with the original market. For everyone else, ordinary copyright rules still apply in full, with none of the goodwill judges have extended to "transformative" AI training.
What actually got decided in 2025
This was the year the theory finally met a judge. Three rulings matter most.
In Bartz v. Anthropic, Judge William Alsup found that training Claude on books Anthropic had legally purchased and digitized was "exceedingly transformative" fair use. But he drew a hard line at the more than seven million books Anthropic pulled from pirate libraries like LibGen to build a permanent internal library. That piracy was not fair use, and it sent Anthropic to a settlement rather than a jury.1 Anthropic agreed to pay $1.5 billion, the largest copyright settlement in U.S. history, working out to roughly $3,000 per book across an estimated 500,000 titles.12
Two days later, in Kadrey v. Meta, Judge Vince Chhabria also found Meta's training use highly transformative and granted summary judgment for Meta, even though Meta had also used shadow-library books. The difference: the authors in that case had not built a record showing market harm. Chhabria was blunt about the limits of his own ruling, writing that in most cases like this, plaintiffs with better evidence "will often win."3
Then there is Thomson Reuters v. Ross Intelligence, decided in February 2025, which cuts the other way entirely. A Delaware court found that training a competing legal-research tool on Westlaw's headnotes was not fair use, because a licensing market for that data plausibly existed. The court stated flatly that it didn't matter whether Thomson Reuters had used the data itself: "the effect on a potential market for AI training data is enough."
Read together, the pattern is: transformative use helps you, pirated sourcing hurts you, and a plausible licensing market for your data hurts you most of all. None of that is a blanket permission slip. The U.S. Copyright Office reached a similar conclusion in its May 2025 report, stating that after weighing all four fair use factors, the analysis "generally favors copyright owners," not AI companies.
The hypocrisy is not subtle
Here is the part that should bother anyone who has ever had a video pulled down or a blog post flagged. Meta has separately used the DMCA, the same law it argues shouldn't restrict AI training, to protect its own leaked Llama model weights. When an early version of Llama was torrented and posted to GitHub in 2023, Meta sent a DMCA takedown demanding it be removed, on the grounds that no one was authorized to reproduce or distribute its property without permission. Meanwhile, in comments to the U.S. Copyright Office, Meta has argued that copyrighted material scraped from across the internet for free should not be protected under copyright when it's used to train Llama.
That is the whole asymmetry in one company's own paperwork. Your copyrighted work is training data. Their model weights are property.
Individual creators do not get the benefit of the doubt, and they rarely get the benefit of a legal team either. Survey data from the Copyright Alliance found that of small creators who discover infringement of their own work online, only about a third actually file a formal DMCA notice, and of those, more than half report the content stayed up anyway even after a valid notice. Meanwhile 94% of the same respondents said they had never once received a takedown notice against them, which tells you who the enforcement machinery is actually built to protect.
Why waiting for the law to catch up is not a plan
Even the judges who ruled in favor of AI companies were careful to say this is not settled law. Chhabria's opinion in the Meta case is explicit that future plaintiffs with a stronger record on market harm could win. The Ross Intelligence ruling shows a completely different outcome is possible when a licensing market exists. And the Copyright Office's own read of the four fair use factors leans toward creators, not toward AI companies.
That means the rules are still being written case by case, largely by whoever can afford to litigate for two years. If you are a company building internal AI tools, or a person feeding proprietary customer data, contracts, or your own writing into someone else's model, you are betting on a legal theory that only holds up when the party using it has a $1.5 billion settlement in the budget.
The safer path: own your data, own your model
The practical lesson isn't about picking a side in the copyright fight. It's about not needing to. If your training data and your models live somewhere you control, the provenance question that has been deciding these cases, licensed versus scraped versus pirated, simply doesn't apply to you the same way, because you're not redistributing anyone else's copyrighted corpus at scale to compete with them.
For internal tools built on your own company's proprietary data, this matters twice over. First, you avoid inheriting legal exposure from a vendor's training pipeline you can't audit. Second, you keep the thing that's actually valuable, your own accumulated data, as an asset instead of quietly handing it to whichever SaaS or model provider you route it through. That's the same argument this publication has been making about renting versus owning your software stack. If you're already thinking about which internal tools are worth building rather than renting, platforms like Remy are built around that same idea: keep the software, and the data flowing through it, as something your team owns outright.
None of this makes the legal picture simple. But it does make the risk asymmetric in your favor for once, instead of against you.
FAQ
Is scraping copyrighted content to train an AI model legal? It depends on where the data came from and what the model competes with. U.S. courts in 2025 found that training on lawfully acquired works can be fair use, but training on pirated copies is not, and training on data where a licensing market exists is also less likely to qualify.
Did Anthropic actually lose its AI copyright case? Partially. Anthropic won on the core question of whether training itself is transformative fair use, but lost on how it acquired millions of the books, which were downloaded from pirate libraries. That piracy claim led to a $1.5 billion settlement rather than a trial.1
Does this mean individuals can freely use AI-scraped content the same way big labs do? No. These rulings turned on facts specific to well-funded defendants with purchased libraries and years of litigation budget. An individual scraping and republishing copyrighted material has none of that cover and faces the same strict infringement rules as always.
What does the U.S. Copyright Office actually recommend? Its May 2025 report concluded the fair use analysis generally favors copyright owners over AI companies, and recommended Congress consider licensing frameworks so creators are compensated when their work trains commercial models.
What's the practical takeaway for a company building AI tools internally? Control your own data and, where possible, your own models. It removes you from the provenance fight altogether and keeps your data as an asset you hold rather than something absorbed into someone else's training pipeline.
It depends on provenance and market effect. Courts found training on lawfully acquired works can be fair use, but training on pirated copies or data with an existing licensing market is not.
Partially. It won on the transformative-use question but lost on using pirated books, leading to a $1.5 billion settlement.
No. These rulings were fact-specific to defendants with purchased libraries and years of litigation budget. Individuals face the same strict copyright rules as before.
Its 2025 report found the fair use analysis generally favors copyright owners and recommended Congress consider licensing frameworks for AI training data.
Control your own training data and models where possible, so you are not inheriting an unaudited vendor's data provenance risk.



