OpenAI, Inc., Copyright Infringement
This MDL consolidates copyright-infringement litigation brought against an artificial-intelligence company and a technology-company co-defendant by a coalition of individual authors, an authors' trade association, and news organizations. Centralized in the Southern District of New York in 2025, the consolidated actions allege that the defendants downloaded and reproduced copyrighted books and journalism to train large language models, and that the resulting AI outputs separately reproduce protected expression from those works. The docket carries 19 pending actions drawn from cases originally filed in multiple federal districts.
What drives resolution risk and timing in this docket is a genuinely unsettled area of copyright law: whether and to what extent using copyrighted works to train a large language model qualifies as fair use, a question with limited controlling appellate precedent at the time this docket was centralized and one being litigated simultaneously across several other AI-training copyright disputes outside this MDL. That legal uncertainty, rather than a factual dispute over what was copied, is the central driver of both duration and outcome range here, since a fair-use ruling in this docket or a related case could materially reshape the litigation landscape for AI-training copyright claims generally.
For anyone tracking how copyright law is adapting to generative AI, this docket is one of the central pending disputes on the question of whether large-scale use of copyrighted text for model training constitutes fair use, and its resolution will likely be read well beyond the specific parties involved. Criterica Intelligence's platform tracks this kind of doctrinally unsettled, precedent-shaping litigation distinctly from disputes governed by well-established legal frameworks, since duration and outcome forecasting for the former depends far more on how courts resolve open legal questions than on case-specific facts.
A coalition of individual authors, an authors' trade association, and news organizations allege that the defendants downloaded and reproduced their copyrighted books and journalism to train large language models, and that the resulting AI outputs separately reproduce protected expression from those works.
Whether large-scale use of copyrighted text to train an AI model qualifies as fair use is a genuinely unsettled legal question with limited controlling precedent, and how it is resolved will likely determine liability more than any factual dispute over what was copied.
No. Similar AI-training copyright disputes are being litigated in parallel outside this MDL, and a fair-use ruling here or in one of those related cases could reshape the legal landscape for the broader category of claims.
Because its duration and outcome depend primarily on how courts resolve an open, precedent-shaping legal question rather than on case-specific facts, a different analytical profile than disputes governed by well-established doctrine.
Statistics shown reflect historical or illustrative model outputs derived from real case data. They are not predictions or guarantees of any individual outcome. Litigation results depend on facts, jurisdiction, judge, and counsel, and vary case by case. Model accuracy is subject to selection effects and changing legal dynamics.