← All posts
case-law

Authors Guild v. OpenAI: The Consolidated Case, Explained

Before there was a New York Times lawsuit, before Anthropic's $1.5 billion settlement, before almost any of the AI copyright litigation getting attention today, there was Authors Guild v. OpenAI. Filed in 2023 by novelists including Mona Awad, Jonathan Franzen, and George R.R. Martin, coordinated with a set of earlier individual suits from authors Paul Tremblay, Michael Chabon, and comedian and writer Sarah Silverman, this was one of the first serious tests of whether training an AI model on copyrighted books without permission is lawful. It's still being litigated, and it's easy to mix up with a different, unrelated case involving a similarly named plaintiff.

Untangling the two Silverman cases

This is worth being precise about, because the names genuinely overlap in a confusing way. Sarah Silverman is a plaintiff in two separate lawsuits against two separate defendants. Silverman v. OpenAI, filed in 2023, is one of the cases now coordinated with Authors Guild v. OpenAI and the Tremblay and Chabon suits, all against OpenAI, over ChatGPT's training data. Silverman v. Meta is a completely different lawsuit, consolidated instead with Kadrey v. Meta, against Meta specifically, over Llama's training data. Same plaintiff, different case, different defendant, different court history, different current status. If you see "Silverman" referenced in coverage of AI copyright litigation, check which defendant is named before assuming which case is being discussed.

Where the case stands

The Tremblay, Chabon, Silverman, and Authors Guild actions against OpenAI are now coordinated as In re: OpenAI, Inc., Copyright Infringement Litigation, proceeding before Judge Sidney Stein in the Southern District of New York, the same court handling the New York Times case, though as a formally separate proceeding.

The most significant recent development came in October 2025, when Judge Stein denied OpenAI's motion to dismiss the consolidated plaintiffs' core infringement claim, allowing the litigation to move forward past the threshold stage where a large share of copyright suits against AI companies have stalled or been narrowed. That ruling didn't decide the merits. It decided that the authors' claims were legally sufficient to proceed into discovery and further litigation, which is a meaningful hurdle cleared but not the same as a finding of infringement.

Since then, the case has moved into active discovery, with public filings from late 2025 into 2026 showing disputes over deleted datasets, privilege assertions around internal OpenAI communications, and the scope of what OpenAI has to disclose about how it built and maintains its training datasets. As of the most recent public filings, there's been no final judgment, settlement, or dismissal, and the litigation remains active with no trial date yet set.

Why this case matters independent of the Times case

It's easy to let the New York Times litigation dominate attention, since it involves a marquee plaintiff and headline-grabbing discovery fights over 20 million chat logs. But Authors Guild v. OpenAI is arguably the more foundational case for a specific reason: it's about books, not news articles, and it's brought by the authors themselves rather than a publisher. The claims center on OpenAI's use of copyrighted novels and other books to train its models, the same basic category of claim that produced Anthropic's $1.5 billion settlement in Bartz. If a similar piracy-and-fair-use analysis eventually applies here, and if OpenAI's book-sourcing practices resemble what discovery revealed about Anthropic's, this case could produce a comparably significant outcome, simply later, because it's moving through discovery more slowly.

The disputes over deleted datasets specifically are worth watching. In several other AI copyright cases, what a defendant did or didn't preserve, and whether records of original training data sources still exist, has turned out to be as consequential as the underlying legal fair-use question. If OpenAI's training data provenance turns out to be murkier or harder to reconstruct than Anthropic's was, that could cut against OpenAI in ways that go beyond the legal merits of the fair-use argument itself.

What this means for AI authors

If you're an author whose books might be part of a large language model's training data, this case, more than almost any other in this roundup, represents your direct stake in this fight. It's brought by novelists, over books, using the same basic legal theory that's already produced a $1.5 billion result against a different company. Whether or not you're one of the named plaintiffs, the outcome here will likely shape how every other AI company handles book-training-data claims going forward.

For your own practice, the deleted-dataset discovery dispute carries a broader lesson worth internalizing: provenance and record-keeping cut both ways. When an AI company can't clearly account for where its training data came from, that becomes a legal liability for the company. When you can't clearly account for how you created and revised your own AI-assisted work, that becomes a liability for you, in a registration, a licensing negotiation, or a dispute over your own authorship. The same principle, that a clean, contemporaneous record beats an unclear or reconstructed one, is running through both sides of this litigation, and it's worth taking seriously in your own creative process regardless of how the industry-level case eventually resolves.

Copyrightable is not a law firm and doesn't provide legal advice. This litigation is active and unresolved. Confirm current status at official court dockets before relying on this summary.

Related reading

Related reading

Want the record behind your own work? See how the methodology scores authorship.

Want a contemporaneous record of how you authored your work?

Try it free