OpenAI, Microsoft, and the New York Times: What a Sealed Memo Tells Us About AI Data Economics

A single sentence from a sealed internal document is circulating as fact — and the gap between what the public can read and what it cannot is the real story.

Case At A Glance — Allegations Only, Nothing Adjudicated

91,692+

Alleged copies of works from the Times, Daily News & CIR in training data — per NYT brief, unadjudicated

93%

Alleged peak drop in NYT click-through rates associated with Copilot — per NYT brief, unadjudicated

Jan 2023

Date of the sealed internal memo from which a single sentence has been quoted in the NYT brief

0

Court findings, rulings, or damages figures — no court has ruled on any claim in this case

What Happened

Unredacted filings in the New York Times’s copyright lawsuit against OpenAI and Microsoft, surfaced by TechCrunch on September 17, 2026, quote a phrase attributed to Brent Hecht, described as Microsoft’s director of applied science: that AI data practices amounted to “the largest theft of labor in human history.” The phrase appears in the Times’s legal brief, which cites a January 2023 internal memo. That memo remains sealed. No reader outside the proceedings can access it, and no court has ruled on any claim in this case.

The same brief — written by the Times as a party to the litigation, in order to advance the Times’s legal position — also alleges that training datasets contained 91,692 or more copies of works from the Times, the Daily News, and the Center for Investigative Reporting; that copyright notices were stripped from training data; and that Microsoft’s Copilot was associated with New York Times click-through rates falling by as much as 93%. Every one of these is an allegation. None has been established by a court. None is stated here as false. OpenAI and Microsoft did not return requests for comment, and no inference is drawn from that.

This is not legal advice. The allegations described in this article are unproven. Nothing here alleges wrongdoing by any person or company. Nothing has been adjudicated, and no court has ruled on any claim in this case.

The key insight: The sentence generating the most circulation is the one the public is least equipped to evaluate — because the document it came from is sealed, the argument surrounding it is unavailable, and a vivid fragment always travels farther than the paragraph explaining its evidentiary status.

The Structural Read

Strip this story of its drama and what remains is a precise problem in information architecture: the unit of evidence in public circulation is a characterisation, not a document. A brief is an argumentative act — an entirely legitimate and ordinary one, performed by both sides in every adversarial proceeding. Choosing which sentence of a long internal document to quote is part of that argument. Nothing here suggests that any party quoted unfairly, misleadingly, or in bad faith. But the consequence for readers is specific and worth naming.

When the underlying document is sealed, the public receives a selection without receiving the thing it was selected from. The same sentence — “the largest theft of labor in human history” — is consistent with an internal objection to a practice, an advocacy for a policy change, a restatement of an external criticism being answered, or a description of an industry rather than any single employer. Each of those is an ordinary thing to find in a researcher’s internal memo. Each produces the identical quotation once the surrounding argument is removed. Nothing here ascribes a motive, position, or state of mind to Brent Hecht. Nothing here says he admitted, confessed, conceded, revealed, or exposed anything.

The structural point is this: an internal critic and an internal admission produce indistinguishable text once context is stripped. Context is the part that carries the meaning. That is precisely why the sealed status of the memo matters more than the sentence itself.

Permission Layer — Structural Observation

“The Permission Layer governing AI data use is not yet settled law — it is actively being constructed through proceedings like this one. What courts ultimately decide about training data will determine which AI business models are structurally viable and which require renegotiation. The public narrative is running years ahead of the legal record.”

The 93% click-through figure deserves its own frame, separated from the legal proceeding. As an allegation, it cannot be adopted here or contradicted here. But the category of claim it represents is a genuine business-model question, independent of this case.

The central economic tension between any AI summariser and any content source is whether the summary substitutes for the source or refers to it. These are structurally opposite relationships. A referral sends an audience onward and can be worth paying for — it is additive to the publisher’s traffic. A substitution satisfies the demand at the point of query and ends the user’s journey before it reaches the source. A click-through rate is one of the few quantities that actually distinguishes between these two dynamics, which is why a figure of that kind sits at the centre of disputes structurally similar to this one. Nothing here says which relationship obtains in this instance, and no other traffic or revenue figure for any publisher appears.

Three Implications

IMPLICATION 1 — The Referral vs. Substitution Question Has No Settled Answer

Every major content publisher is now exposed to the same structural uncertainty: does an AI product that surfaces their content drive audience to them or satisfy demand in place of them? No court, regulator, or platform has produced a definitive answer. The click-through rate is the empirical test — and publishers who cannot measure it cannot negotiate from evidence.

IMPLICATION 2 — Sealed Records and Selective Quotation Are Standard Litigation Tools, Not Anomalies

The dynamic on display here — a vivid fragment from an unavailable document entering public circulation with velocity — is not a failure of any party, journalist, or platform. It is how adversarial proceedings interact with media distribution. Analysts and readers who treat quoted phrases from sealed documents as established facts will systematically misread the evidentiary record in any major tech lawsuit that follows.

IMPLICATION 3 — The Permission Layer Governing AI Training Data Remains Structurally Unsettled

Until courts rule — not allege, not brief, but rule — on whether and how training data use must be licensed, compensated, or constrained, AI companies are building on a Permission Layer whose final shape is unknown. That uncertainty is itself a business risk, priced differently by every player in the stack, and it will not resolve through public narrative alone.

Business Engineer Framework

The Permission Layer

In the Business Engineer Map of AI, the Permission Layer is the regulatory and legal stratum that controls which AI products can ship, at what scale, and on what data. This case is Permission Layer in real time: the outcome will draw a boundary — or decline to — around how AI training interacts with copyrighted content at commercial scale. Understanding the layer beneath the headlines is how you model the actual business risk.

Explore the Map of AI →

The Bottom Line

The most important thing about “the largest theft of labor in human history” is not whether it is true, false, damning, or exculpatory — no one outside the proceedings can assess that, because the document it came from is sealed and no court has ruled on anything. What matters, for anyone building or investing in AI, is that the legal architecture governing training data is being actively constructed in proceedings like this one, and the public narrative is running years ahead of the actual record. A sentence that travels at headline speed, stripped of a sealed context, is a signal about how information moves — not a finding about what happened.


This article is not legal advice. The allegations described are unproven. Nothing here alleges wrongdoing by any person or company. No court has ruled on any claim in this case. Sources: TechCrunch, September 17, 2026.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This is not legal advice. The claims described above are unproven allegations made by one party in ongoing litigation, and nothing above alleges wrongdoing by any person or company. No court has ruled on any of them. The January 2023 memo remains sealed, and the quoted phrase reaches the public through a brief filed by the opposing party; what argument that memo was making is not publicly known, and nothing above characterises Brent Hecht as admitting, confessing or revealing anything, or ascribes any view to him. Nothing above suggests that any party has quoted unfairly, misleadingly or in bad faith — selecting quotations is the ordinary work of an adversarial brief on both sides. OpenAI and Microsoft did not return requests for comment, and no inference is drawn from that above. The observations that a selection cannot be assessed without the document it came from, that substitution and referral are different economic relationships, and that circulation strips qualifiers before it strips claims are general properties. They are not findings about this case, and nothing above says which party is right or how the matter should be resolved. Nothing is predicted.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA