← All posts
case-law

News Corp and Dow Jones v. Perplexity: When "Just Search" Meets Copyright

News Corp and Dow Jones filed suit against Perplexity AI in October 2024, and for a while it looked like just another entry on the growing list of publisher-versus-AI-company lawsuits. Then, on August 21, 2025, Judge Katherine Polk Failla denied Perplexity's motion to dismiss on jurisdictional grounds, and the case moved into discovery. That ruling matters more than the filing date, because it means this one is actually going to be tested on the merits, not thrown out on a technicality.

What Perplexity does differently

Every other major AI copyright case you've probably read about, the New York Times suit, Getty v. Stability, the various music-label cases, targets a foundation model: a system trained on huge datasets that later generates new text, images, or audio. Perplexity isn't that. It's a search and answer engine. You ask it a question, and it goes out, retrieves information (often from the same paywalled articles News Corp and Dow Jones publish), and hands you back a synthesized answer, sometimes with quoted passages baked in.

That's a different kind of product, and News Corp and Dow Jones are testing a different kind of legal theory against it. Their complaint alleges Perplexity fed entire paywalled stories from the Wall Street Journal, the New York Post, and other News Corp properties into its indexing and retrieval system without a license, and that its outputs sometimes reproduce verbatim paragraphs from those articles mixed together with fabricated additions, sentences the underlying article never actually said, presented with the same authority as the real reporting. The amended complaint, filed in December 2024, expanded the claims to include false designation of origin and trademark dilution alongside copyright infringement, essentially arguing that Perplexity is putting its own invented content next to real reporting and letting readers assume both came from the same trustworthy source.

Perplexity's defense

Perplexity's position is that what it does is fair use, no different in kind from what a search engine or a research tool has always done: index publicly available and licensed content, then summarize and cite it for users. Google has built a trillion-dollar business on a version of that argument, and search engines have generally won fair use fights over indexing in the past. Perplexity's lawyers are leaning on that precedent hard.

The problem for Perplexity is that "we just index and summarize" gets harder to defend the closer your output gets to reproducing the original text. A search engine sending you to the original article is a different act than a chatbot handing you the article's paragraphs directly, inside its own interface, without you ever clicking through to the source. News Corp and Dow Jones are betting that a judge and eventually a jury will see the difference, especially once discovery surfaces exactly how much of the original text Perplexity's system stores, retrieves, and reproduces versus how much it genuinely rewrites.

Why the jurisdictional ruling matters

Failla's August 2025 ruling didn't touch the merits of the copyright or trademark claims at all. It rejected Perplexity's argument that the case shouldn't be heard in the venue News Corp and Dow Jones chose, or that the company lacked sufficient contacts with the jurisdiction to be sued there. That's a procedural win, but it's the win that keeps the whole case alive. Cases that get dismissed on jurisdictional grounds never reach a jury, and they rarely produce the kind of discovery record that shapes how other courts think about similar disputes. This one now will.

Bloomberg Law's coverage of the suit described what's at stake in blunt terms: publishers see AI answer engines as an "existential threat" to the traffic and subscription revenue their businesses depend on. If a reader gets the Wall Street Journal's reporting summarized inside a chatbot response, with no need to visit wsj.com, subscribe, or see an ad, the publisher's entire business model is bypassed at the exact moment the content is consumed. That's a sharper version of the concern publishers have raised against search engines for years, just faster and more direct.

What this means for AI authors

If you build or use retrieval-based AI tools, meaning anything that pulls in outside content at query time rather than relying purely on what a model learned during training, this case is the one to watch, not the training-data suits. The legal question here isn't "was this content used to train a model." It's "does this product reproduce protected expression when it answers a question," which is a much more direct and much easier claim for a plaintiff to prove than trying to reverse-engineer what happened inside a training run months or years earlier.

The "we just search and summarize" defense sounds airtight until you look closely at what your product actually outputs. If your tool quotes source material closely enough that a reader gets the substance of the original without needing the original, you're exposed to the same theory News Corp and Dow Jones are running against Perplexity, regardless of how small or unknown your company is compared to Perplexity. Scale doesn't create the legal risk here. Reproduction does.

And if you're a writer or publisher wondering whether your own paywalled or licensed content might be getting scraped and reproduced by an answer engine somewhere, the practical lesson from this case is the same one that runs through nearly every suit in this space: what matters in court is evidence of what was actually taken and how it was actually used, not a general sense that "AI is probably using my stuff." Keep records of what you published and when. That's the foundation any claim like this eventually has to stand on.

Related reading

Want the record behind your own work? See how the methodology scores authorship.

Want a contemporaneous record of how you authored your work?

Try it free