The New York Times v. OpenAI and Microsoft, Explained
Twenty million chat logs. That's the number at the center of the discovery fight in The New York Times v. OpenAI and Microsoft, and it tells you something about the scale this case has grown to, more than two years after the Times first filed. This is the most closely watched active AI copyright case in the country, and as of August 2026 it's still years from a trial date.
What happened
The New York Times sued OpenAI and Microsoft in the Southern District of New York, alleging that ChatGPT and Microsoft's Copilot products can reproduce near-verbatim excerpts of Times journalism, sometimes long enough and specific enough to function as a substitute for reading the article itself. Where the Bartz and Kadrey cases focus mainly on what happened during training, the Times's complaint leans heavily on outputs, what the model actually produces when you ask it about Times reporting, which makes it a more direct test of whether these tools can function as a free alternative to the underlying journalism they were trained on.
The case has been in discovery for an extended period, and the discovery fights themselves have become a significant part of the story. Judge Stein affirmed an order requiring OpenAI to produce roughly 20 million de-identified user chat logs on January 5, 2026, a scale of discovery that's unusual even by the standards of complex commercial litigation and gives some sense of how seriously the court is treating the question of what ChatGPT actually outputs in response to real user prompts.
The June 2026 pivot toward Microsoft
In June 2026, the Times amended its complaint in a way worth understanding clearly: it dropped its contributory infringement claims against OpenAI specifically and refocused its secondary liability theory on Microsoft, alleging Microsoft built dedicated infrastructure, reportedly a purpose-built supercomputer with more than 285,000 CPU cores and over 10,000 GPUs, specifically to support OpenAI's model training and operation.
This wasn't a random strategic choice. It followed a US Supreme Court ruling in a different case, Cox Communications, decided March 25, 2026, that reshaped how secondary liability claims (holding a company responsible for someone else's infringement, rather than its own direct infringement) have to be pled going forward. The Times's amendment is widely read as an adaptation to that new legal landscape, aiming its secondary liability theory at the party, Microsoft, that arguably had the clearest and most direct role in building and providing the infrastructure the alleged infringement ran on.
Separately, in July 2026, the Times filed a sanctions motion alleging OpenAI misrepresented its search capabilities to the court for more than two years, a filing that, if it goes anywhere, adds a credibility dimension to the litigation on top of the substantive copyright claims.
Why the Cox Communications ruling changed the Times's strategy
The Supreme Court decision that triggered the June amendment deserves its own explanation, because it's reshaping secondary-liability arguments across every AI copyright case, not just this one. In Cox Communications, Inc. v. Sony Music Entertainment, decided March 25, 2026, the Court unanimously narrowed what it takes to hold a company contributorily liable for someone else's infringement. The prior standard let plaintiffs argue a company was liable if it merely knew infringement was happening on its platform and kept providing service anyway. The Court rejected that. Contributory liability now requires actual intent, shown either by affirmatively inducing the infringement or by providing a product or service specifically tailored to enable it. Mere knowledge, even repeated knowledge, isn't enough anymore.
That ruling made the Times's original contributory-infringement theory against OpenAI considerably harder to win, since it had leaned partly on OpenAI's awareness that its models could reproduce protected content. Rather than fight an uphill battle on a weakened legal theory, the Times dropped that specific claim against OpenAI and its related trademark claims, and pivoted its secondary-liability theory entirely toward Microsoft, where it believes it has a stronger "tailored to infringe" argument: that Microsoft didn't just provide general cloud infrastructure to a customer, but built a custom supercomputing system, according to the amended complaint more than 285,000 CPU cores and over 10,000 GPUs, specifically to enable large-scale training on copyrighted material at the volume OpenAI needed. That's a meaningfully different claim than "you hosted the servers," and it's a direct response to the higher bar Cox Communications just set.
The sanctions fight over search capabilities
Days after that amendment, the litigation got another twist. On July 9, 2026, the Times and a coalition of other publishers, including the New York Daily News, the Center for Investigative Reporting, The Intercept, and Ziff Davis, filed a motion asking the court to sanction OpenAI, alleging the company had repeatedly told the court it lacked the technical ability to search its own training datasets and ChatGPT output logs for the publishers' copyrighted content, only for a later deposition of an OpenAI corporate witness to reveal the company had already built exactly those search tools, and had used them internally before the litigation began. The publishers are asking for real consequences: barring OpenAI from relying on its 20-million-log sample, a court finding that the full logs would have shown "substantial and systematic" reproduction of the publishers' work, attorney's fees, and jury instructions reflecting evidence destruction. That motion was pending as of this writing.
Why this case matters more than most
Every major AI company is watching this one closely, and not just because the Times is a sophisticated, well-resourced plaintiff. It's because the case tests the output side of the AI copyright question directly, can a model reproduce your specific protected expression closely enough to matter, rather than the training side that most other cases focus on. A ruling against OpenAI and Microsoft here wouldn't just be about training data practices. It would go to whether the fundamental function of these products, generating responses to user queries, can itself constitute infringement when the response reproduces someone else's work too closely.
There's no trial date yet, and given the scale of discovery already underway, one realistically isn't expected before late 2026 or into 2027 at the earliest. This is a case to watch unfold over years, not months.
What this means for AI authors
If you're a writer, journalist, or content creator, this case is the clearest test yet of a question that affects you directly: can an AI tool reproduce your specific writing closely enough, in response to a user's question, that it functions as a substitute for people actually reading your work. That's a different and in some ways more concrete harm than "my work was used to train a model," and it's worth tracking this case's outcome if your business depends on people actually visiting your content rather than getting it secondhand through a chatbot.
If you use AI tools yourself to draft or research content, this case is a reminder that AI-generated text can carry real infringement risk when it closely tracks someone else's specific expression, independent of any question about your own authorship of the surrounding work. A model can produce output that's too close to a copyrighted source, and that risk sits with whoever publishes the output, regardless of how much editing or original work surrounds it. Being able to show what you wrote yourself versus what a tool generated, and how closely you checked generated text against source material, is exactly the kind of record that matters if this question ever comes up in your own work.
Related reading
Related reading
- Every AI Copyright Lawsuit Worth Knowing: The Complete Guide
- Zarya of the Dawn: What the Copyright Office Actually Kept
- Theatre D'opera Spatial: 624 Prompts Weren't Enough
Want the record behind your own work? See how the methodology scores authorship.
Want a contemporaneous record of how you authored your work?
Try it free