Issue #284

Microsoft's Own Memo: LLMs Are Eating Their Supply Chain

Unsealed court filings show Microsoft and OpenAI knew the risk to publishers long before this lawsuit.

BusinessMicrosoft's Own Memo: LLMs Are Eating Their Supply Chain

The Sentence That Will Outlast “The Greatest Theft of Labor in History”

On the 17th (local time), a batch of previously sealed documents from The New York Times’s copyright lawsuit against OpenAI and Microsoft became public. Nearly every headline covering the news ran nearly the same line: “the greatest theft of labor in human history,” a phrase Brent Hecht, Microsoft’s Director of Applied Science, reportedly wrote in an internal memo in January 2023.

It’s a headline-worthy sentence. But I lingered longer on a different line from the same author: large language models are a product that destroys their own supply chain. Microsoft’s internal documents apparently included a passage to this effect, too — that it’s exceedingly rare for a final product to threaten the economic foundation of its own key suppliers, and that this is exactly the situation the company had created in its LLM business.

Theft is the language of ethics — a word to be argued over in court. Supply chain is the language of business. Today I want to talk about the latter: what the company knew internally, what it confirmed with numbers, and why, knowing all that, it still didn’t move to protect its suppliers.

What the Unsealed Documents Say

Let me first lay out what came to light. The lawsuit was filed by The New York Times in December 2023, and 11 other news organizations, including the New York Daily News, have since joined it. Judge Sydney Stein of the U.S. District Court for the Southern District of New York is reviewing a motion for summary judgment1, and documents are surfacing piece by piece as part of that process.

What’s been disclosed this time breaks down into roughly three strands.

The first is scale. According to the plaintiffs’ filings, OpenAI’s intermediate training dataset alone contained more than 91,692 copies of articles from the plaintiff publications, and one dataset drawn from Common Crawl2 included more than 2,000,000 nytimes.com documents by itself. The plaintiffs also allege circumstantial evidence that the company circumvented paywalls and stripped copyright notices from training data — including an exchange in which, after an OpenAI researcher reported a “hack” for bypassing The New York Times’s paywall, President Greg Brockman reportedly replied, “oh nice.”

The second is internal concern. It wasn’t just Hecht’s memo. Nick Turley, who led ChatGPT’s development, wrote in a June 2023 memo that AI posed an existential threat to news publishers, and in February 2024 noted that as AI products got better, they would become increasingly substitutive. Back in 2020, Jack Clark, then OpenAI’s policy director, sent a memo to Sam Altman and Brockman warning that the company was building systems that would replace the labor of the people who shape society’s culture. Clark later left the company to co-found Anthropic.

The third is executive testimony. Microsoft CEO Satya Nadella said in a deposition that anyone wanting to use content behind a paywall needed to obtain a license for it. He also said that if he had known OpenAI had scraped and trained on paywalled information, he would have exercised the company’s right to demand that the model be retrained.

Two caveats are worth keeping in mind, though. Much of what’s come out so far consists not of primary evidence but of excerpts quoted in The New York Times’s own filings — meaning these are sentences cut out of their surrounding context, since the original documents remain sealed. And Microsoft has pushed back on the idea that Hecht’s memo reflects the company’s position, stating in its own filings that he wasn’t a decision-maker but a researcher hired specifically to offer a contrarian perspective. OpenAI did not respond to a request for comment.

The Word Heavier Than “Theft” Is “Substitute”

The headline was theft, but I think there’s a different word that’s going to carry weight in the trial. That word is substitute.

OpenAI and Microsoft’s defense rests on fair use3 under U.S. copyright law. Their argument is that using articles for training transforms the originals into something sufficiently new — that it’s not a substitute that damages the original’s market value. One of the factors courts weigh in fair use analysis is exactly this: what effect does the use have on the market for the original work.

But in this filing, the word “substitute” — used internally at these companies — keeps coming up. Someone wrote that the AI product being sued over is, in fact, mostly a substitute, and appended “end of story.” In February 2023, an OpenAI engineer wrote to colleagues that no matter how prominently you displayed a link, users simply wouldn’t click it. And Microsoft’s own data, cited in the brief, showed that Copilot’s answer engine drove click-through rates to nytimes.com down by as much as 93% compared to standard Bing search.

We’ve touched on this link issue in our newsletter before. When Google rolled out its preferred-source button for publishers, we cited outside research showing that AI summaries cut link clicks by roughly half. What’s different this time is the source. It’s not an outside research firm — it’s internal data from the company that built the product, and an engineer’s own words. The very claim that “adding source links sends traffic back” wasn’t something the builders themselves believed.

Google’s Preferred Source Button Doesn’t Tell You Your Follower Count · Issue 210 · INLEVEL9A breakdown of what Google’s newly disclosed preferred-source embed button actually does, and why publishers still can’t see how many readers designated their site or how much traffic it sends.inlevel9.com

None of this means we can predict how the lawsuit ends. So far, U.S. courts have mostly leaned toward treating AI training as fair use, and earlier this month the Trump administration filed a brief supporting OpenAI’s position. Excerpts stripped of context can always be read differently at trial. Still, the fact that a company arguing “it’s transformation, not substitution” had someone inside write “it’s a substitute, period” — that’s now on the record, regardless of how the verdict lands.

Buyers Who Starve Their Suppliers to Death

Let’s get back to the supply chain story.

If you’re running a business, a key supplier collapsing isn’t first an ethics problem — it’s a survival problem. When I wrote about SanDisk last August, I saw exactly the opposite scene play out. Customers building AI data centers signed contracts with a NAND supplier running longer than 4 years, backing them with $16.5 billion in guarantees that would pay out immediately if the deals fell through. If the parts stop coming, their own business stops — so the buyers put money into the supplier’s survival first.

SanDisk’s 4-Year Contracts, Backed by $16.5 Billion in Guarantees · Issue No. 183 · INLEVEL98 customers demanded supply contracts running 4+ years and $16.5 billion in financial guarantees upfrontinlevel9.com

What makes Hecht’s document interesting is that it applies this exact same logic to content. He called the phenomenon of Copilot draining clicks from publishers a “doom loop.” If publishers weaken, good writing declines; if good writing declines, model performance declines too; and the whole web gets worse together. This wasn’t an ethical appeal — it was a warning about the raw-material supply for our own product.

But unlike semiconductors, buyers didn’t hold onto their content suppliers. The document doesn’t directly answer why. From here on, this is my own interpretation, but I think two things overlapped.

One is timing. If NAND supply is cut today, servers stop today. Articles, by contrast, are already trained on decades’ worth of text, so the cost of a supplier collapsing doesn’t come due for years. The other is the number of suppliers. NAND suppliers can be counted on one hand, but there are countless places writing content. Even if one collapses, substitutes seem readily available — for now. When these two conditions overlap, even if you know the supply chain is collapsing, the incentive to pay for it today weakens.

Meanwhile, the price on the supplier side started getting set through lawsuits and individual contracts. Some publishers, like News Corp, signed major deals with OpenAI; some authors, like me, licensed a single book for AI training and received a few hundred dollars. But this looks less like a deliberately designed supply chain and more like patches hastily applied while watching the whole thing fall apart.

Oswarld’s Lens

There’s one passage in this document I kept staring at longer than the rest: Microsoft’s explanation of Hecht. He was hired to voice a different perspective — he wasn’t a decision-maker.

supply chainAs a legal defense, that holds up. But as a description of organizational design, it’s strange. They hired someone specifically to raise objections, that person did exactly that, and now the company is saying the outcome wasn’t its position. It reads as an admission that the warning existed but never reached the seat where decisions get made. Nadella’s testimony has the same shape. Saying he would have demanded retraining if he’d known, flipped around, means information that could have reached his desk simply never did. Of course, testimony is testimony — it may also just be self-protective language.

Set alongside Clark’s 2020 memo, the picture comes into sharper focus. The warning wasn’t a one-off, and it wasn’t confined to a single company. Across 2020, 2023, and 2024, essentially the same warning was written down multiple times, inside two different companies. And still, the direction never changed.

Companies rolling out AI these days tend to create certain roles: ethics committees, red teams, “responsible AI” leads. I think these roles are necessary. But what this case shows is that creating a seat for dissent and building a pathway for that dissent to actually reach the decision are two entirely different things. Without that pathway, the memo only gets read properly years later — in a courtroom, and as evidence for the other side.

So I think the better first question to ask any organization is this: when a memo saying “this will wreck our supply chain” gets written, whose desk does it stop at?

Closing

This document doesn’t settle the verdict. The excerpts are missing context, and courts still lean toward fair use. But one thing remains: the sentences written inside the AI company pinned down the problem more precisely than the criticism coming from outside ever did. Not the word “theft,” but the word “supply chain.”


💬 Does your organization have a path where internal dissent can actually reach the people who make decisions? If you’ve seen that path change a direction, or watched it die before arriving anywhere, tell us in the comments.

📨 If you have a colleague weighing an AI rollout or a content business, pass this along.


Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Related issues worth reading


Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Summary judgment: A procedure by which a judge rules on a case based purely on questions of law — without a full trial or jury — when there’s no genuine dispute over the facts. Because both sides aggressively submit whatever evidence favors them at this stage, previously sealed documents often surface here. ↩

  2. Common Crawl: A nonprofit archive that scrapes web pages at massive scale and makes them freely available. Many AI labs use it as a starting point for training data. ↩

  3. Fair use: A principle in US copyright law that allows copyrighted work to be used without permission in certain cases — criticism, reporting, parody, and the like. Courts weigh four factors together: the purpose and character of the use, the nature of the original work, the amount used, and the effect on the market for the original. ↩