The AI Training Loophole Just Closed

ideko

A federal appeals court just told the AI industry something it really didn’t want to hear: you can’t scrape someone’s copyrighted work, use it to train a competitor, and call it fair use. And honestly? It’s about time.

A federal appeals court just told the AI industry something it really didn’t want to hear: you can’t scrape someone’s copyrighted work, use it to train a competitor, and call it fair use. And honestly? It’s about time.

What Actually Happened Here

The Third Circuit ruled that Ross Intelligence – a legal AI startup – straight-up infringed Thomson Reuters’ copyrights when it trained its system on Westlaw headnotes. For anyone not steeped in legal research tools, Westlaw is basically the Google of case law, and those headnotes? They’re these carefully written summaries that Westlaw editors create to help lawyers find the good stuff in court opinions fast.

Ross wanted to build an AI legal research tool. Cool idea, actually. The problem was how they went about it.

The AI Training Loophole Just Closed

They partnered with some outfit called LegalEase Solutions to create training materials that connected legal questions with relevant passages from court opinions. Sounds fine until you realize those training materials were built using Westlaw’s headnotes – the exact thing Thomson Reuters pays editors to write and considers its competitive advantage.

So Ross essentially used Westlaw’s own editorial work to build a product designed to compete with Westlaw. The audacity is kind of impressive, not gonna lie.

Why This Matters (And It Really Does)

This case cuts right to the heart of the biggest question hanging over the AI boom: when is training on copyrighted material actually legal?

Tech companies have been operating under this assumption – this hope, really – that training AI models on copyrighted content is protected fair use. The logic goes something like: we’re not copying your work, we’re learning from it. It’s transformative. It’s research. It’s progress.

And look, there’s been some truth to that argument in certain contexts. But here’s the thing the Third Circuit spotted immediately: Ross wasn’t doing some abstract research project. They were building a direct competitor using their rival’s copyrighted editorial content.

The AI Training Loophole Just Closed

That’s not transformation. That’s just regular old copyright infringement with extra steps and a neural network.

The Fair Use Fantasy

Fair use has four factors courts are supposed to weigh, and Ross basically failed on all of them. The purpose? Commercial, and designed to compete. The nature of the work? Creative editorial summaries, not raw public domain case law. The amount used? Enough to train a working system. The market effect? They literally built a Westlaw alternative using Westlaw’s content.

When you lay it out like that, it’s almost embarrassing how weak the fair use argument was.

“The court found Ross Intelligence used Westlaw’s headnotes to build a direct competitor”

The AI Industry Is Sweating

This ruling lands right in the middle of multiple ongoing fights about AI training data. OpenAI is getting sued by the New York Times. Stability AI got hit with lawsuits from artists. GitHub Copilot faced a class action over code licensing. That's especially notable now that OpenAI's safety team was gutted overnight, raising questions about who's left to think through these risks.

All of them have versions of the same defense: training is fair use, we’re not reproducing the exact content, our output is transformative.

But the Third Circuit just narrowed that loophole considerably. If you’re using copyrighted training data to build something that competes in the same market as the copyright holder, you’re going to have a really hard time claiming fair use. And you should.

What This Actually Means

I mean, it doesn’t kill AI development. Let’s be clear about that. Companies can still license training data (imagine that – actually paying for the content you use). They can train on public domain material. They can use content they create or own. They can probably still get away with some genuinely transformative uses.

What they can’t do – or at least, what just got a lot riskier – is hoover up copyrighted content from a competitor or content provider, train a model on it, and launch a competing product. Which, when you think about it, seems pretty reasonable?

The AI industry has been sprinting forward on the assumption that copyright law wouldn’t catch up or wouldn’t apply to them the way it applies to everyone else. This ruling is a reality check. Training data isn’t free just because it’s digital and you’re calling your product “AI.” Meta's new AI Muse shows exactly that sprint in action, churning out ad content at a scale that makes licensing look like an afterthought.

Thomson Reuters won this round because Ross got too greedy and too obvious about it. But the implications stretch way beyond legal research tools. Every AI company that trained on scraped copyrighted content to build commercial products just got a lot more nervous about their litigation exposure. It's part of why some investors, like Michael Burry, are openly rooting for a crash to kill AI's IPOs before the hype outruns the legal risk.

And honestly? Maybe that’s not such a bad thing.

Share:

Emily Carter

Emily Carter is a seasoned tech journalist who writes about innovation, startups, and the future of digital transformation. With a background in computer science and a passion for storytelling, Emily makes complex tech topics accessible to everyday readers while keeping an eye on what’s next in AI, cybersecurity, and consumer tech.

Related Posts