← All posts
case-law

Every AI Copyright Lawsuit Worth Knowing: The Complete Guide

Nineteen lawsuits. Three countries. A $1.5 billion settlement, a $3 billion suit still pending, a Supreme Court cert denial, and a German court that already reached a final answer on a question American courts are still years from resolving. If you've tried to follow AI copyright litigation by reading headlines as they break, you've probably ended up with a jumbled, contradictory picture: one case says AI training is fine, another says it isn't, a third settled for a fortune, and somehow they're all supposedly about the same legal question.

They're not, and that's the whole point of this guide. AI copyright litigation isn't one fight. It's two, running in parallel, testing completely different legal questions, and mixing them up is the single most common way people misread this entire area of law.

The two-track split

Track one: authorship and registrability. These cases ask whether a specific work, already made, can be copyrighted at all, and if so, by whom. Thaler v. Perlmutter, Zarya of the Dawn, Theatre D'Opera Spatial, and A Single Piece of American Cheese all sit here. Nobody in these cases is arguing about whether AI companies did something wrong by training their models. The question is narrower and more personal: did a human do enough of the creative work on this specific piece to own it. This is the track that matters most directly to you if you're a writer, artist, developer, or musician using AI tools in your own process.

Track two: training-data infringement. These cases ask whether the AI companies themselves did something unlawful in building their models, whether that's how they acquired training data, what their models output, or both. Bartz v. Anthropic, the New York Times suit, Getty v. Stability, Disney v. Midjourney, Kadrey v. Meta, the Concord and BMG suits, Andersen v. Stability, Authors Guild v. OpenAI, GEMA v. OpenAI, News Corp and Dow Jones v. Perplexity, Doe v. GitHub, and the News/Media Alliance coalition suit all sit here. This is the track that matters most if you're a rights holder wondering whether your existing work was used without permission, or if you're trying to understand the legal exposure baked into the AI tools you use every day.

There's also a third category worth naming even though it's not strictly copyright litigation: cases like Sage v. Lovo and Standing v. TikTok, where AI cloned someone's voice without genuine consent. Voice isn't copyrightable, so these run on right-of-publicity, fraud, and contract theories instead. They're covered in this roundup because the underlying harm, an AI system reproducing something a person made without authorization, rhymes with everything else here, even though the legal doctrine is different.

Confusing the two tracks produces exactly the kind of bad takes that circulate after every ruling: "a court said AI training is legal" (usually an overread of a narrow, fact-specific ruling in one training-data case) getting cited as if it settles the completely separate question of whether your own AI-assisted novel can be copyrighted. It doesn't. These are different legal doctrines, different plaintiffs, different courts, and in most cases, different outcomes so far.

Track one: the authorship cases

Thaler v. Perlmutter is the floor everything else sits on. Stephen Thaler tried to register a work with an AI system listed as the sole author. The courts said no, a work needs a human author, full stop. The Supreme Court denied certiorari on March 2, 2026, making this final. Nothing in the rest of this pillar means anything if there's no human author at all.

Zarya of the Dawn drew the real line above that floor. Kristina Kashtanova's text and arrangement stayed protected. The individual Midjourney-generated images she selected and placed did not, because she didn't control the specific pixels the model produced.

Theatre D'Opera Spatial tested how much prompting is enough, and answered: even 624 detailed prompts don't convert instruction into authorship of the underlying image.

A Single Piece of American Cheese is the flip side, a case where human selection and arrangement of AI-assisted elements cleared the bar and got a real registration.

Together, these four cases give you the actual test the Copyright Office applies: not how much AI was involved, but whether a human controlled the specific expressive choices in the final work. Our Copyrightability hub and USCO Standard hub go deep on what that means in practice.

Track two: the training-data cases

This is where the real money and the real corporate exposure sit, and it splits again into two sub-patterns worth understanding.

The piracy pattern. Bartz v. Anthropic established the defining distinction in this entire track: training on legally acquired books can be fair use, but acquiring training data through piracy, torrenting from shadow libraries instead of paying for licenses, is not excused just because the material was later put to a transformative use. That reasoning produced a $1.5 billion settlement, the largest copyright settlement in US history, and it's now the template other plaintiffs are following. Concord Music Group, UMG, and BMG's suits against Anthropic apply the identical theory to song lyrics and sheet music, alleging the same shadow-library sourcing, now in a different industry, worth more than $3 billion combined. Kadrey v. Meta has piracy in its background facts too, and even though Meta won summary judgment on the training claim, the judge said so explicitly because of a gap in the plaintiffs' market-harm evidence, not because training on pirated books is categorically fine, and distribution-based claims tied to the piracy itself are still alive.

The output and reproduction pattern. The New York Times v. OpenAI and Microsoft leans on what ChatGPT actually outputs when prompted about Times journalism, not just what it was trained on, and the case has grown into one of the largest discovery fights in the history of this kind of litigation. GEMA v. OpenAI reached a decisive answer on a closely related theory in Germany already: when a model has memorized specific lyrics closely enough to reproduce them on demand, that reproduction infringes, regardless of how the training itself is characterized. Disney, NBCUniversal, DreamWorks, and Warner Bros. Discovery's suit against Midjourney is the character-specific version of the same idea: can you prompt a tool for Darth Vader or Elsa and get a recognizable copy back, and does the company let you.

The provenance and trademark pattern. Getty Images v. Stability AI is the case that shows where the traction actually is right now when a straightforward copyright claim struggles. Getty lost its core copyright argument in the UK largely on a territorial technicality, and its US case saw its DMCA claim dismissed too. What survived and moved forward in both courts was narrower and different: trademark and unfair-competition claims tied to Stable Diffusion generating images that carried Getty's own watermark, misrepresenting where the image actually came from. That's not a ruling about whether training was lawful. It's a ruling that misleading the public about a work's origin is its own, more durable legal problem.

The long-haul cases. Andersen v. Stability AI is the earliest-filed case in this entire wave, from January 2023, and it's still working toward class certification and summary judgment, both scheduled for hearing in November 2026, with a trial date set for April 2027 if claims survive that far. It's the reminder that visual-art claims face a harder evidentiary road than text-based ones, because proving a specific image is substantially similar to specific training data is a much more technical fight than proving a chatbot reproduced a paragraph. Authors Guild v. OpenAI (not to be confused with Kadrey v. Meta, despite a shared plaintiff name) is moving faster on its output-based infringement theory, which survived a motion to dismiss in October 2025 and continues in discovery.

The retrieval and search pattern. Every case above targets a foundation model trained on a dataset. News Corp and Dow Jones v. Perplexity targets something different: an AI answer engine that retrieves and reproduces paywalled news at query time, not just during training. Judge Katherine Polk Failla rejected Perplexity's jurisdictional challenge in August 2025, so the case is now moving to the merits of whether "we just index and summarize" holds up when the output reproduces verbatim paragraphs. It's the clearest test yet of whether search-style AI products face the same exposure as the training-focused labs.

The attribution and DMCA pattern. Doe v. GitHub is the case built around code rather than prose, testing whether Copilot has to reproduce your open-source code exactly to trigger liability, or whether stripping attribution from a close paraphrase violates the DMCA's copyright-management-information protections too. The Ninth Circuit heard argument on that question in February 2026. The same DMCA theory, stripping bylines and copyright notices before training, shows up again in the News/Media Alliance coalition's suit against OpenAI and Microsoft, filed June 2026 on behalf of nine major publishers including Condé Nast, The Atlantic, and The Guardian. A related suit by overlapping publishers against Cohere Inc., a far less famous enterprise AI vendor, shows this exposure isn't limited to the handful of AI labs that make headlines.

What's explicitly out of scope here

Elon Musk's suit against OpenAI over the company's nonprofit-to-for-profit conversion is not a copyright case. It's a dispute about corporate governance and mission, resolved by jury verdict in May 2026 and then dismissed as time-barred. It doesn't belong in this roundup and citing it as an AI copyright precedent is a category error.

Sage v. Lovo and Standing v. TikTok sit in a different gray area. They're included here because they're part of the same wave of AI-versus-creator litigation, but they aren't copyright cases either. Voice actors deceived into recording samples that were then cloned into commercial products, and a voice actress whose voice allegedly powered TikTok's text-to-speech feature without consent, are suing on right-of-publicity, fraud, and breach-of-contract theories, because your voice, unlike your writing or your music, isn't something copyright protects. Worth knowing if you're a performer, and worth not confusing with the copyright claims that make up the rest of this list.

The cross-case lessons for AI authors

Pull back from the individual cases and a few patterns hold across nearly all nineteen.

Acquisition method matters more than use. Across Bartz, Concord, BMG, and the background facts in Kadrey, courts and settling parties have consistently treated "how was the training data obtained" as a more tractable, more decisive question than "was training itself transformative." If you're a rights holder, the piracy-based theory is currently the strongest, cleanest legal path available, and it's the one producing real money. If you're an AI company or a heavy AI tool user, this is exactly why provenance and licensed sourcing matter more than most people assume.

A single ruling almost never settles the underlying question. Kadrey v. Meta is cited constantly as "AI training is legal." The judge who wrote that opinion said, in the same opinion, that he expects most future cases with better market-harm evidence to come out the other way. Treat every headline ruling in this space as a data point about one case's specific record, not a verdict on the whole field.

Provenance and trademark theories are quietly outperforming pure copyright theories. Getty's watermark claim survived in two countries while its core copyright claim struggled in both. That's not a coincidence. Proving what a work's origin actually is, and whether an output falsely represents that origin, is a more settled, more winnable body of law than the still-unresolved question of whether AI training itself infringes.

Memory, alone, loses to a record. Every case in this roundup eventually turns on evidence: who acquired what, when, from where, and what a model actually does with it. The Copyright Office side of this (Track One) turns on the same thing at the individual level: whether you can show, not just claim, what you actually authored. Contemporaneous records consistently beat after-the-fact reconstruction, in courtrooms and in registration applications alike.

This is a multi-year process, not a news cycle. Andersen has been active since January 2023 and won't see a trial before 2027 at the earliest, if it gets that far. The NYT case is realistically years from resolution. Don't expect a single ruling to answer "is AI copyright settled yet." It won't be, for a long time, and treating any one case as the final word is how you end up acting on legal advice that expires the next time a different court rules differently on a different record.

What to actually do with this

If you're a creator using AI tools, the training-data litigation mostly isn't something you can control. What you can control is your own side of Track One: keeping a real record of the creative decisions you make, so your own authorship claim doesn't depend on memory or reconstruction if it's ever challenged. If you're a rights holder wondering whether your work shows up in a training set somewhere, the piracy-acquisition theory running through Bartz, Concord, and BMG is currently your strongest opening, and it depends on documentation just as much as the authorship side does.

Either direction, the same principle holds across every single case in this roundup: the party with real, specific, contemporaneous evidence has the stronger position, whichever side of the courtroom they're standing on.

Related articles in this pillar

Want a contemporaneous record of how you authored your work?

Try it free