Microsoft, OpenAI Sued by Papers Seeking Removal From Models

Every claim here is an allegation in the plaintiffs’ complaint. This publication read the complaint and the court docket and verified none of it independently.

Six publishing companies, led by Emmerich Newspapers, Inc., filed a complaint in the U.S. District Court for the Southern District of Mississippi against Microsoft and ten OpenAI entities. The complaint, stamped “Filed 10/02/26” (case 3:26-cv-00791), alleges copyright infringement and the removal of copyright management information in building the GPT models.

Among the relief it asks for is “An injunction under 17 U.S.C. § 503(b) requiring Defendants to remove all copies of Registered Works from all GPT or other LLM models and training sets.” Everything below is an allegation in that document, pleaded on information and belief. This publication read the complaint and the CourtListener docket and verified none of it independently.

Who Is Suing Whom

The plaintiffs are Emmerich Newspapers, Inc.; Ojai Media, LLC; Coopwood Publishing Group, Inc.; Coopwood Magazine, Inc.; Coopwood Media Group, Inc.; and Coopwood Newspapers, Inc. The complaint calls Emmerich “one of the largest privately owned newspaper chains in Mississippi, with additional newspapers in Louisiana and Arkansas.”

It says Ojai owns the Ojai Valley News in California, founded in 1891. It says the Coopwood companies, based in Cleveland, Mississippi, publish titles including the Mississippi Business Journal and Delta Magazine.

The caption names Microsoft Corporation and ten OpenAI entities: OpenAI, Inc.; OpenAI LP; OpenAI GP, LLC; OpenAI, LLC; OpenAI OpCo, LLC; OpenAI Global, LLC; OAI Corporation, LLC; OpenAI Holdings, LLC; OpenAI Foundation; and OpenAI Group PBC. The complaint says the OpenAI defendants “consist of a web of interrelated entities.”

CourtListener lists the case as assigned to Henry Travillion Wingate. Its docket, last updated on 5 October 2026 at 2:32 p.m., shows three entries, all dated 5 October: the complaint, an issued summons, and a “Notice – (Copyright, Patent, Trademark).” The complaint is signed by lawyers at Barrett Law Group, P.A.; Cuneo Gilbert Flannery & LaDuca, LLP; and Wilson Carroll, PLLC.

Our count of the numbered paragraphs in the complaint, grouped under its own headings. The count says nothing
Our count of the numbered paragraphs in the complaint, grouped under its own headings. The count says nothing about the strength of any claim.

Business Pill · TAKING ONE THING BACK OUT

A one-minute explainer of the idea behind this story: unlearning. It teaches the concept, not this story’s figures.

The key insight: The relief the plaintiffs ask for follows from where the complaint says copies sit: not only in datasets and outputs but in the models themselves. The complaint says model parameters “encode retrievable copies,” and its prayer asks for removal of copies from “all GPT or other LLM models and training sets.” It describes no method for that removal, and its support for the plaintiffs’ own articles is two dataset analyses it does not detail, further allegations about dataset contents made without a described analysis, and testing in other cases.

How the Complaint Says the Articles Reached the Models

The complaint says defendants “systematically and secretly crawled Plaintiffs’ websites, including content behind paywalls and other access restrictions,” and copied the articles onto their own servers. It says the process was “repeated over and over and over again” as the models were updated.

It names the datasets it says were used. For GPT-2, it cites OpenAI’s own description of a dataset called WebText, built from “the text contents of 45 million links posted by users of the ‘Reddit’ social network.” For GPT-3, it says the mix included Common Crawl, WebText2, Books1, Books2 and Wikipedia content.

The complaint calls Common Crawl a “copy of the Internet” made available by a 501(c)(3) organization of the same name. It describes two analyses of datasets. Paragraph 68 says: “A straightforward analysis of the Common Crawl dataset clearly shows that tens of thousands of Plaintiffs’ copyrighted articles and tens of millions of tokens worth of Plaintiffs’ content likely appeared in the Common Crawl training dataset Defendants used to train GPT-3.”

Paragraph 98 says an analysis “performed by a technologist consultant employed by Plaintiffs’ counsel” found that OpenWebText, “an open-source approximation of the WebText dataset,” contained “thousands of records totaling millions of tokens of text from Plaintiffs’ websites.”

Neither paragraph describes how its analysis was done or lists the articles found. The docket lists one attachment to the complaint, a civil cover sheet.

What It Says About Outputs and Memorization

The complaint says models trained this way show what researchers call memorization, and that this phenomenon “shows that LLM parameters encode retrievable copies of many of those training works.”

It adds that products using retrieval, such as ChatGPT Search and Deep Research, “retrieve content from Plaintiffs’ websites in real time and incorporate that content, oftentimes verbatim or in close paraphrase, into synthesized responses delivered to users.”

For what the models produce, the complaint points to other cases. It says testing in “substantially identical copyright infringement actions” brought by The New York Times, the New York Daily News and other publishers showed GPT-based models reproducing “substantial portions of published news articles when prompted appropriately.” It says, on information and belief, that OpenAI has since disabled public access to the versions of ChatGPT tested in those cases.

For its own plaintiffs, the complaint says on information and belief that those earlier versions “would likewise generate near-verbatim copies of Plaintiffs’ copyrighted works if prompted to do so.” It says further information about the output of the plaintiffs’ content “is solely in OpenAI’s possession, custody, and control.” The text sets out no example output from a plaintiff’s article.

The Copyright Management Information Allegation

The complaint says defendants’ systems stripped “author credits, publication names, copyright notices, and terms of use information” from the articles. It quotes OpenAI’s description of how it extracted text from web pages for WebText: “a combination of the Dragnet and Newspaper content extractors.”

It alleges, on information and belief, that OpenAI used those extractors “to intentionally remove CMI.” It also says the C4 dataset, which it describes as a snapshot of Common Crawl, “contains the full text of online news articles by Plaintiffs devoid of the publication name, article title, subtitle, byline, date, copyright notice, and terms of use links.”

Willfulness and Notice

The complaint cites OpenAI’s written evidence to a House of Lords committee, dated 5 December 2023, for the statement that “it would be impossible to train today’s leading AI models without using copyrighted materials.” It says the plaintiffs put defendants on notice by placing copyright notices and links to their terms of service on every page of their websites that defendants copied.

It describes the defendants’ conduct as “systematic and willful theft” of the plaintiffs’ articles. That is the plaintiffs’ language, not this publication’s.

The Three Counts

Count I alleges direct copyright infringement under 17 U.S.C. § 501. The complaint says Microsoft and the OpenAI defendants directly infringed by building training datasets, by storing and processing them on Microsoft’s supercomputing platform, and, on information and belief, by storing models that “have memorized” the works. It says OpenAI infringed by disseminating output through ChatGPT, and Microsoft by doing so through Copilot.

Count II alleges vicarious infringement. It says Microsoft “controlled, directed, and profited from” infringement by the OpenAI defendants and “had the ability to control the infringing activity” but failed to do so. Paragraph 152 names OpenAI, Inc.; OpenAI LP; OAI Corporation, LLC; OpenAI Holdings, LLC; OpenAI Global, LLC; and Microsoft as vicariously liable.

Count III alleges removal of copyright management information under 17 U.S.C. § 1202(b). Its paragraphs refer to “OpenAI,” which the complaint defines as the OpenAI entities, and not to Microsoft.

On Microsoft’s role, the complaint says the infrastructure included “a dedicated supercomputer that Microsoft built and operated for OpenAI’s exclusive use.” It says, on information and belief, that Microsoft “was intimately involved in every step of OpenAI’s LLM training.”

The Relief: Removal From Models and Training Sets

The prayer for relief asks for statutory damages, “including willful infringement damages,” plus compensatory damages, restitution and disgorgement. No dollar amount appears in the prayer. It asks for costs and attorney’s fees, and the complaint demands a jury trial.

It asks for a permanent injunction against “the unlawful, unfair, and infringing conduct alleged herein.” The opening paragraph describes the action as one for “preliminary and permanent injunctive relief”; the prayer, as read, lists a permanent injunction.

It also asks for “An injunction under 17 U.S.C. § 503(b) requiring Defendants to remove all copies of Registered Works from all GPT or other LLM models and training sets.” That request follows the complaint’s theory of copying: it says copies sit in the training datasets, in the models’ parameters, and in the outputs. The prayer does not say how removal from a trained model would be carried out.

The complaint also cites an OpenAI post of 31 March 2026 for a post-money valuation of $852 billion, says OpenAI generates $2 billion in monthly revenue, and says OpenAI filed confidentially for an IPO in June 2026, citing The New York Times. These are the plaintiffs’ pleaded figures. This publication did not check them.

The Structural Read

The complaint pleads a chain with four links: articles on the plaintiffs’ websites, copies in datasets such as Common Crawl and WebText, training of the GPT models, and outputs through ChatGPT and Copilot. At each link it alleges copying. The support it points to differs by link: for the second, two dataset analyses plus further allegations about dataset contents made without a described analysis; for the fourth, testing in other cases plus information and belief.

The three counts divide the defendants differently. Count I is pleaded against Microsoft and the OpenAI entities, Count II pleads a control theory for Microsoft and, in paragraphs 150-152, vicarious liability for five OpenAI entities as well, and Count III, the copyright-management-information count, is pleaded against “OpenAI” as the complaint defines it.

The removal request sits at the third link. By asking for removal “from all GPT or other LLM models and training sets,” the prayer applies the complaint’s copying theory to the models themselves, not only to the data and the outputs. The prayer does not say how removal from a trained model would be done.

The complaint — prayer for relief, item iii

“An injunction under 17 U.S.C. § 503(b) requiring Defendants to remove all copies of Registered Works from all GPT or other LLM models and training sets”

Three Implications

THE REQUEST REACHES THE MODELS The prayer pairs damages with an order to remove copies from models and training sets. The complaint says model parameters encode “retrievable copies of many of those training works”; the removal request applies to models as well as to training sets.

SUPPORT IS PLEADED, NOT SHOWN For the plaintiffs’ own articles the complaint describes two dataset analyses without describing their method, makes further allegations about dataset contents without describing an analysis, and relies on testing in other cases for outputs. It says on information and belief that earlier model versions would “likewise generate” copies, and it sets out no example output.

WHAT REMAINS UNKNOWN No response from Microsoft or OpenAI was located, and the docket lists no ruling. The complaint’s text does not list the registered works, and this publication did not read the notice filed with it or any source it cites.

What Is Not Established

Every statement above is an allegation in the plaintiffs’ complaint. The court has not ruled; the docket as last updated lists no ruling. This publication did not locate a response from Microsoft or OpenAI in the searches it ran on 5 October.

The complaint does not describe the method behind its two dataset analyses, and its text does not list the registered works at issue. This publication did not read the notice filed with the complaint, any registration record, or any source the complaint cites.

This publication did not contact the plaintiffs, their lawyers, Microsoft or OpenAI, and offers no view on the merits or on whether the relief sought is available.

Business Engineer Framework

The Map of AI — Where Training Data and Models Sit in the Stack

The Map of AI places more than 200 companies across nine layers of the stack, from silicon to application. Training data and the models built on it sit in the middle layers, and the Map shows who occupies the layers around them.

Read the Map of AI →

The Bottom Line

Six publishers allege in a complaint stamped as filed on 2 October 2026 that Microsoft and ten OpenAI entities copied their articles into training data, removed copyright management information and reproduce the works through ChatGPT and Copilot. Among other relief, they ask for an injunction requiring removal of all copies of their registered works from GPT or other LLM models and training sets. All of it is allegation; the court has not ruled, and this publication did not locate a response from either company. This publication verified none of it independently.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This piece rests on the complaint filed in Emmerich Newspapers, Inc. et al. v. Microsoft Corp. et al. (S.D. Miss. 3:26-cv-00791) and the CourtListener docket page, both read on 5 October 2026. This publication did not contact the plaintiffs, their lawyers, Microsoft or OpenAI, and did not read any document the complaint cites. Nothing above predicts anything, and nothing here is legal or investment advice.

Sources: courtlistener.com · Emmerich Newspapers, Inc. et al. v. Microsoft Corp. et al., S.D. Miss. 3:26-cv-00791, Complaint (Doc. 1, filed 10 · CourtListener docket 74915508 (last updated Oct. 5, 2026, 2:32 p.m.)

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA