AI Copyright Infringement: What Actually Counts
AI Copyright Infringement: What Actually Counts
People throw the phrase "AI copyright infringement" around to mean at least three different things, and mixing them up leads to bad conclusions. There's infringement in what AI companies fed into their models. There's infringement in what an AI tool spits back out. And there's a totally separate question, not about infringement at all, about whether AI-assisted output can be copyrighted in the first place. Let's untangle them.
Training-data infringement: did the AI company have the right to use the material
This is the biggest category of active litigation, and it's about the input side of the equation. When a company builds a generative AI model, it trains that model on enormous datasets, often scraped from the open internet, that include copyrighted books, articles, images, and music. The legal question is whether using that copyrighted material to train a model, without a license, is infringement or whether it falls under fair use.
The results so far are mixed, and that mixed record is the headline. In the Bartz case against Anthropic, a federal judge found that training on lawfully purchased books could be fair use, but that Anthropic's use of pirated copies to build its training library was not protected, and the case ultimately settled for $1.5 billion, the largest copyright recovery in the history of these disputes. The New York Times' case against OpenAI and Microsoft is still working through discovery and remains one of the most closely watched suits in the field. Getty Images sued Stability AI in both UK and US courts over training on its licensed photo library, with different outcomes taking shape in each jurisdiction. Music publishers went after Suno and Udio directly, arguing that training generative music models on copyrighted recordings without a license is straightforward infringement, not some novel AI exception.
None of these cases turn on a single, unified rule yet. Judges are looking at specifics: was the training data acquired legally, was the use transformative, does the output compete with the market for the original work. We track the full docket, case by case, in our complete guide to every AI copyright lawsuit worth knowing.
Output infringement: did the AI generate something that copies existing work
This is a different question from training-data infringement, even though people often conflate the two. Here, the issue isn't how the model was built. It's whether a specific thing the model generated is substantially similar to a specific existing copyrighted work, the same test that's applied to human-made infringement claims.
Generative AI models can and do reproduce recognizable elements of their training data, sometimes verbatim or near-verbatim. A widely discussed example: media companies allege that certain prompts can cause an AI system to output passages that closely track a copyrighted article. Disney and Warner Bros. sued Midjourney over its ability to generate images of copyrighted characters on request, arguing that a tool designed to output recognizable trademarked and copyrighted characters is a different problem than a tool that merely learned general artistic style from a broad dataset. That's an output-side claim: the argument isn't about what Midjourney was trained on in the abstract, it's about what specific images it produces and how closely they track protected characters.
If you're a user of these tools, this is the category where you have the most personal exposure. If you generate an image or passage of text that closely mirrors an existing copyrighted work, and you publish or sell it, you can be on the hook for infringement even if you didn't intend to copy anything, the same way you would be if you'd copied it by hand.
The question that isn't about infringement at all: is your AI-assisted work copyrightable
Here's where the confusion really sets in. A lot of people ask "is AI content copyright infringement" when what they actually mean is "can I get a copyright on something I made with AI." Those are opposite directions. Infringement is about whether you took something you shouldn't have. Copyrightability is about whether what you made is original enough, and human-authored enough, to protect in the first place.
The Copyright Office has been clear that AI-generated material, standing alone with no human creative input, isn't eligible for protection at all, following the reasoning in Thaler v. Perlmutter. But a human-authored arrangement, edit, or selection built on top of AI output can be. That's a copyrightability question, governed by how much creative control a person exercised, and it has nothing to do with whether anyone infringed anyone else's rights. We go deep on that distinction in our piece on what counts as enough human authorship.
A concrete way to sort it
If you're trying to figure out which bucket a specific situation falls into, ask these in order:
Is the question about how the AI model itself was built and trained? That's a training-data infringement question, and it's a fight between rights holders and AI companies, not something an individual user is typically exposed to.
Is the question about whether a specific output resembles a specific existing copyrighted work too closely? That's output infringement, and it's the one category where an ordinary user of these tools carries real personal risk. Publishing an AI output that's substantially similar to something copyrighted can expose you the same way republishing someone else's work without permission would.
Is the question about whether you can register a copyright on something you made using AI? That's not an infringement question at all. That's copyrightability, and it turns entirely on how much human creative control you exercised over the final expression.
Keeping these three questions separate clears up most of the confusion in this space. Most people asking about "AI copyright infringement" online are actually asking a copyrightability question, and the answer to that one has nothing to do with whether anybody infringed anything.
Related reading
- Can AI-Generated Content Be Copyrighted?
- Is AI-Generated Code Copyrightable?
- Can You Copyright AI Art?
- Every AI Copyright Lawsuit Worth Knowing: The Complete Guide
Want the record behind your own work? See capture your authorship record from the first prompt.
Want a contemporaneous record of how you authored your work?
Try it free