Criterica Intelligence — production models trained on real court records, not synthetic data
Intellectual Property — MDL No. 3143

OpenAI, Inc., Copyright Infringement

U.S. District Court for the Southern District of New York

This MDL consolidates copyright-infringement litigation brought against an artificial-intelligence company and a technology-company co-defendant by a coalition of individual authors, an authors' trade association, and news organizations. Centralized in the Southern District of New York in 2025, the consolidated actions allege that the defendants downloaded and reproduced copyrighted books and journalism to train large language models, and that the resulting AI outputs separately reproduce protected expression from those works. The docket carries 19 pending actions drawn from cases originally filed in multiple federal districts.

What drives resolution risk and timing in this docket is a genuinely unsettled area of copyright law: whether and to what extent using copyrighted works to train a large language model qualifies as fair use, a question with limited controlling appellate precedent at the time this docket was centralized and one being litigated simultaneously across several other AI-training copyright disputes outside this MDL. That legal uncertainty, rather than a factual dispute over what was copied, is the central driver of both duration and outcome range here, since a fair-use ruling in this docket or a related case could materially reshape the litigation landscape for AI-training copyright claims generally.

For anyone tracking how copyright law is adapting to generative AI, this docket is one of the central pending disputes on the question of whether large-scale use of copyrighted text for model training constitutes fair use, and its resolution will likely be read well beyond the specific parties involved. Criterica Intelligence's platform tracks this kind of doctrinally unsettled, precedent-shaping litigation distinctly from disputes governed by well-established legal frameworks, since duration and outcome forecasting for the former depends far more on how courts resolve open legal questions than on case-specific facts.

Frequently Asked Questions
What is alleged in the OpenAI copyright MDL?

A coalition of individual authors, an authors' trade association, and news organizations allege that the defendants downloaded and reproduced their copyrighted books and journalism to train large language models, and that the resulting AI outputs separately reproduce protected expression from those works.

Why is the fair-use question so central to this docket's outcome?

Whether large-scale use of copyrighted text to train an AI model qualifies as fair use is a genuinely unsettled legal question with limited controlling precedent, and how it is resolved will likely determine liability more than any factual dispute over what was copied.

Is this dispute unique to this one MDL?

No. Similar AI-training copyright disputes are being litigated in parallel outside this MDL, and a fair-use ruling here or in one of those related cases could reshape the legal landscape for the broader category of claims.

Why does Criterica Intelligence treat this docket differently from other IP MDLs?

Because its duration and outcome depend primarily on how courts resolve an open, precedent-shaping legal question rather than on case-specific facts, a different analytical profile than disputes governed by well-established doctrine.

Statistics shown reflect historical or illustrative model outputs derived from real case data. They are not predictions or guarantees of any individual outcome. Litigation results depend on facts, jurisdiction, judge, and counsel, and vary case by case. Model accuracy is subject to selection effects and changing legal dynamics.

← All Pending MDLsFunding brief on Criterica Capital →