Software Ownership

We Need to Talk About Who Owns the AI Training Data

Big AI labs scraped the internet, got sued, and mostly won or settled their way clear. If you tried the same thing with their content, you would not be so lucky. That asymmetry is the whole story.

Minimal ink and crimson illustration of a corporate hand scooping copyrighted documents into a glowing AI server while a single locked personal folder sits untouched on a desk

The short answer

AI scraping is fair use only sometimes, and mostly for the companies with the lawyers to prove it in court. U.S. courts ruled in 2025 that training a generative model on lawfully acquired copyrighted works can be fair use, but training on pirated copies is not, and the outcome depends heavily on how the data was obtained and whether it competes with the original market. For everyone else, ordinary copyright rules still apply in full, with none of the goodwill judges have extended to "transformative" AI training.

What actually got decided in 2025

This was the year the theory finally met a judge. Three rulings matter most.

In Bartz v. Anthropic, Judge William Alsup found that training Claude on books Anthropic had legally purchased and digitized was "exceedingly transformative" fair use. But he drew a hard line at the more than seven million books Anthropic pulled from pirate libraries like LibGen to build a permanent internal library. That piracy was not fair use, and it sent Anthropic to a settlement rather than a jury.1 Anthropic agreed to pay $1.5 billion, the largest copyright settlement in U.S. history, working out to roughly $3,000 per book across an estimated 500,000 titles.12

Two days later, in Kadrey v. Meta, Judge Vince Chhabria also found Meta's training use highly transformative and granted summary judgment for Meta, even though Meta had also used shadow-library books. The difference: the authors in that case had not built a record showing market harm. Chhabria was blunt about the limits of his own ruling, writing that in most cases like this, plaintiffs with better evidence "will often win."3

Then there is Thomson Reuters v. Ross Intelligence, decided in February 2025, which cuts the other way entirely. A Delaware court found that training a competing legal-research tool on Westlaw's headnotes was not fair use, because a licensing market for that data plausibly existed. The court stated flatly that it didn't matter whether Thomson Reuters had used the data itself: "the effect on a potential market for AI training data is enough."

Read together, the pattern is: transformative use helps you, pirated sourcing hurts you, and a plausible licensing market for your data hurts you most of all. None of that is a blanket permission slip. The U.S. Copyright Office reached a similar conclusion in its May 2025 report, stating that after weighing all four fair use factors, the analysis "generally favors copyright owners," not AI companies.

The hypocrisy is not subtle

Here is the part that should bother anyone who has ever had a video pulled down or a blog post flagged. Meta has separately used the DMCA, the same law it argues shouldn't restrict AI training, to protect its own leaked Llama model weights. When an early version of Llama was torrented and posted to GitHub in 2023, Meta sent a DMCA takedown demanding it be removed, on the grounds that no one was authorized to reproduce or distribute its property without permission. Meanwhile, in comments to the U.S. Copyright Office, Meta has argued that copyrighted material scraped from across the internet for free should not be protected under copyright when it's used to train Llama.

That is the whole asymmetry in one company's own paperwork. Your copyrighted work is training data. Their model weights are property.

Individual creators do not get the benefit of the doubt, and they rarely get the benefit of a legal team either. Survey data from the Copyright Alliance found that of small creators who discover infringement of their own work online, only about a third actually file a formal DMCA notice, and of those, more than half report the content stayed up anyway even after a valid notice. Meanwhile 94% of the same respondents said they had never once received a takedown notice against them, which tells you who the enforcement machinery is actually built to protect.

Why waiting for the law to catch up is not a plan

Even the judges who ruled in favor of AI companies were careful to say this is not settled law. Chhabria's opinion in the Meta case is explicit that future plaintiffs with a stronger record on market harm could win. The Ross Intelligence ruling shows a completely different outcome is possible when a licensing market exists. And the Copyright Office's own read of the four fair use factors leans toward creators, not toward AI companies.

That means the rules are still being written case by case, largely by whoever can afford to litigate for two years. If you are a company building internal AI tools, or a person feeding proprietary customer data, contracts, or your own writing into someone else's model, you are betting on a legal theory that only holds up when the party using it has a $1.5 billion settlement in the budget.

The safer path: own your data, own your model

The practical lesson isn't about picking a side in the copyright fight. It's about not needing to. If your training data and your models live somewhere you control, the provenance question that has been deciding these cases, licensed versus scraped versus pirated, simply doesn't apply to you the same way, because you're not redistributing anyone else's copyrighted corpus at scale to compete with them.

For internal tools built on your own company's proprietary data, this matters twice over. First, you avoid inheriting legal exposure from a vendor's training pipeline you can't audit. Second, you keep the thing that's actually valuable, your own accumulated data, as an asset instead of quietly handing it to whichever SaaS or model provider you route it through. That's the same argument this publication has been making about renting versus owning your software stack. If you're already thinking about which internal tools are worth building rather than renting, platforms like Remy are built around that same idea: keep the software, and the data flowing through it, as something your team owns outright.

None of this makes the legal picture simple. But it does make the risk asymmetric in your favor for once, instead of against you.

FAQ

Is scraping copyrighted content to train an AI model legal? It depends on where the data came from and what the model competes with. U.S. courts in 2025 found that training on lawfully acquired works can be fair use, but training on pirated copies is not, and training on data where a licensing market exists is also less likely to qualify.

Did Anthropic actually lose its AI copyright case? Partially. Anthropic won on the core question of whether training itself is transformative fair use, but lost on how it acquired millions of the books, which were downloaded from pirate libraries. That piracy claim led to a $1.5 billion settlement rather than a trial.1

Does this mean individuals can freely use AI-scraped content the same way big labs do? No. These rulings turned on facts specific to well-funded defendants with purchased libraries and years of litigation budget. An individual scraping and republishing copyrighted material has none of that cover and faces the same strict infringement rules as always.

What does the U.S. Copyright Office actually recommend? Its May 2025 report concluded the fair use analysis generally favors copyright owners over AI companies, and recommended Congress consider licensing frameworks so creators are compensated when their work trains commercial models.

What's the practical takeaway for a company building AI tools internally? Control your own data and, where possible, your own models. It removes you from the provenance fight altogether and keeps your data as an asset you hold rather than something absorbed into someone else's training pipeline.

Figure 1
How the 2025 AI training cases were decided
Fair use finding
1Bartz v. Anthropic (purchased books)0Bartz v. Anthropic (pirated books)1Kadrey v. Meta0Thomson Reuters v. Ross
Based on 2025 summary judgment rulings; Bartz was a split decision covering two different data sources.
Source: Remy analysis
Frequently asked
Is scraping copyrighted content to train an AI model legal?

It depends on provenance and market effect. Courts found training on lawfully acquired works can be fair use, but training on pirated copies or data with an existing licensing market is not.

Did Anthropic actually lose its AI copyright case?

Partially. It won on the transformative-use question but lost on using pirated books, leading to a $1.5 billion settlement.

Can individuals rely on the same fair use arguments as big AI labs?

No. These rulings were fact-specific to defendants with purchased libraries and years of litigation budget. Individuals face the same strict copyright rules as before.

What does the Copyright Office recommend?

Its 2025 report found the fair use analysis generally favors copyright owners and recommended Congress consider licensing frameworks for AI training data.

What should companies do to reduce their legal exposure?

Control your own training data and models where possible, so you are not inheriting an unaudited vendor's data provenance risk.

Sources
  1. 1.Anthropic to pay authors $1.5 billion to settle lawsuit over pirated books — Associated Press
  2. 2.Anthropic settles with authors in first-of-its-kind AI copyright deal — NPR
  3. 3.Fair Use and AI Training: Two Recent Decisions Highlight Uncertainty — Skadden
Portrait of Marcus Bello
Marcus Bello
Build vs Buy
Marcus writes about when teams should build their own tools instead of buying.
© 2026 The Official Remy BlogDrafted by AI authors, reviewed by human editors.