Issue #199

My Book's AI Training Royalty: $300–656 Per Copy

When my book sold for AI training at $300–656 a copy, I realized the real shift is in how content gets priced, not what it's worth.

BusinessMy Book's AI Training Royalty: $300–656 Per Copy

My Book’s AI Training Royalty Came Out to $300–656 a Copy

A little while ago, a slightly strange payment landed in my account. It was a copyright fee for AI training use of a book I’d written, and it came out to $300–656 per copy. At current exchange rates, that’s roughly ₩900,000 (~$650) at most. My publisher brokered the deal, and the sale went through a US company called Mentat (Mentat).

My first reaction to the number wasn’t “cheap” or “expensive.” It was closer to not knowing what, exactly, I was being paid for.

Then, on August 12, I saw a Wall Street Journal report. Apple is in talks with news publishers to fix Siri, and the terms are unfamiliar. Rather than a flat fee, they’re proposing to pay each time the content is actually used.

What’s happening right now isn’t a question of whether creative work’s price is going up or down. The basis for pricing is shifting from the fact that a book was fed into training data to how many times that book actually gets referenced. Music is the industry that went through this same transition first, and we already know how it turned out.

Until Now, Books Were Priced One at a Time

To see where $656 falls, it helps to line up the price tags that have gone public over the past two years.

The most frequently cited benchmark is the HarperCollins–Microsoft deal. Disclosed in November 2024, its terms are fairly specific: $5,000 per book, split 50-50 between author and publisher, so the author’s share is $2,500. The term runs 3 years, and it caps the model’s output so it can’t reproduce more than 200 consecutive words of the original text or more than 5% of the whole book. It was opt-in — included only if the author agreed.

The media-sector numbers are a different order of magnitude. In May 2024, News Corp’s deal with OpenAI was reported at up to $250 million over 5 years. That’s a flat, company-to-company arrangement.

The third benchmark isn’t a deal at all — it’s damages. Anthropic’s $1.5 billion settlement with authors received final approval on July 21, 2026, covering roughly 500,000 works at about $3,000 per work. What matters here is that this isn’t an arm’s-length market price; it’s a figure assigned to unauthorized use.

656Set alongside these three, my $656 finds its place. Comparing author shares alone, it’s a little over a quarter of the HarperCollins figure, and about a fifth of the per-work rate in the Anthropic settlement.

The instinct here is to jump straight to “Korean authors got shortchanged,” but I think that reading misses something important. Language, genre, length, term, and exclusivity all differ, which makes simple comparison difficult. The HarperCollins deal was English-language nonfiction limited to 3 years; my case runs on different terms entirely.

What catches my attention isn’t the amount but the fact that all three price things the same way: so much per book, so much per work. In other words, the price is set the moment a book enters the training dataset — nothing more. However often that book gets cited afterward, or never referenced at all, the amount stays the same.

This approach actually works fairly well for creators. You can predict what you’ll be paid, in advance.

Apple Proposed a Pay-As-You-Go Model

Now let’s go back to this report.

According to WSJ, Apple has approached multiple news organizations in recent months and is proposing multi-year contracts. The budget under discussion runs to 9 figures—at least $100 million. But the payment structure is variable. It’s built so Apple pays in proportion to how much a partner’s content actually gets used.

WSJ specifically flagged this point. The standard contract between major AI companies and news organizations bundles a guaranteed flat fee with broad access rights—and what Apple proposed isn’t that standard.

Apple isn’t the only one. The trend is already moving in the same direction across multiple players.

Microsoft is piloting a Publisher Content Marketplace (PCM). As of February 2026, AP, Business Insider, Condé Nast, Hearst, USA Today, and Vox Media have joined, and the pilot advertises usage-based reporting that pays “according to value delivered.”

Perplexity has built a $42.5 million distribution pool for Comet Plus. The structure gives 80% of subscription revenue to publishers and keeps 20%, with payouts tied to three usage signals: direct visits via the browser, citations in search answers, and use by the AI assistant to complete a task.

Cloudflare is doing the same thing at the infrastructure layer. The payment rails for AI agents and the access toll fees covered in the last issue are examples of this.

Here we need to separate training from grounding.

What Apple needs for Siri isn’t so much model training as grounding1—pulling in the latest news and current information on the fly to answer a query. Training happens once and is done; grounding happens fresh every time a question comes in. That’s why you can count how many times something was referenced.

So pay-per-use is a natural contract for news. New articles appear every day, and each one gets referenced anew every time.

The problem is that once this contract structure becomes the standard, it gets carried over into other content contracts too. What happens when content like books, papers, and lecture notes—material that, once learned, is rarely referenced again—gets pulled into the same structure? The value something contributes to training isn’t measured by reference count, but the settlement system only counts reference count.

Mentat’s business structure shows this shift has already begun. The company brokers content from 38 publishers and more than 21,000 books, and it presents publisher revenue in two forms: a one-time payment for dataset licensing, and recurring revenue based on citation frequency. The $656 I received was the former—a one-time payment. Which means the recurring-revenue-by-citation-frequency line item is already built into the contract terms.

Music and libraries went through this transition first

The music industry went through this same transition earlier. So instead of predicting what happens, we can just look at the results.

As the recording industry moved from sales to streaming, it changed its settlement standard from album sales to per-song play counts. Then, on March 18, 2024, Spotify announced a policy: tracks with fewer than 1,000 plays over the trailing 12 months would no longer earn recording royalties.

Spotify’s explanation is, in its own way, reasonable. Tracks below that threshold account for roughly 0.5% of the total royalty pool, and per track that works out to an average of $0.02 a month — about ₩30 (~$0.02). But since intermediary distributors and banks set minimum withdrawal amounts between $2 and $50, with fees running $1 to $20, that money disappears into fees and minimum-payout thresholds before it ever reaches the artist. The logic is: better to pool it and redirect it to artists who are actually active.

That logic is hard to argue with. But look at the outcome: usage-based pricing created a band where payouts approach zero, and Spotify simply excluded that entire band from payment. And the songs sitting in that band are the overwhelming majority. If 99.5% of all plays come from tracks that clear the threshold, flip that around and it means the vast majority of songs were splitting just 0.5% of plays. In other words, songs in the long tail2 have almost no way to stay eligible for payment under a pure usage-based system — their payouts simply fall below the fees and minimum-withdrawal thresholds we just discussed.

But this experiment is actually much older. And it started out in a far gentler form.

Public Lending Right3 — a scheme that compensates authors based on how many times their books are borrowed from libraries — began in Denmark in 1946. Today, 33 countries run some version of it; the UK adopted it in 1979. According to World Intellectual Property Organization data, about 24,000 creators in the UK receive payments under the scheme every year.

The UK scheme, as described in the WIPO interview cited in this piece, caps payouts at £6,600 per year. Why bother with a cap at all? Because if you pay strictly per loan, a handful of bestselling authors would sweep up nearly all the funds. So the system pays based on loan counts, but caps the amount to correct for money concentrating in too few hands.

Here’s the summary. Pricing by usage is not a new experiment. And every place that’s tried it has, without exception, discovered that a pure usage-based model doesn’t actually work — and has bolted on some kind of corrective mechanism. Spotify drew a floor and cut low-play songs out of payment entirely; the UK’s PLR drew a ceiling and capped how much high-circulation authors could collect. The directions are opposite, but both share the same conclusion: neither settles purely by usage alone.

Right now, AI content contracts still don’t have that corrective mechanism.

Oswarld’s Lens

I’ve worked on several projects that involved shifting flat-rate subscriptions to usage-based pricing while building GTM strategy. Every single time, I confirmed the same thing. A shift to metered pricing looks like a pricing adjustment, but it’s actually a shift in who gets to decide what counts as “usage.”

In a flat-rate contract, the seller negotiates over the amount. Once you move to metered pricing, the object of negotiation shifts from the amount to the method of measurement. What counts as usage, who counts it, who owns the counting system — and the answer to all three questions is almost always held by the buyer.

When Perplexity says it has defined three usage signals, it’s Perplexity that defined those three. When Microsoft says it will pay “based on value delivered,” it’s Microsoft that calculates the value. When Apple says it will pay when content is used, only Apple knows whether it was actually used.

Working with data, I’ve developed a habit: before looking at any metric, I first check who owns the system that aggregates that number. Whoever owns the aggregation ends up holding the pricing power in that market. In metered content deals, usage aggregation happens entirely inside the AI company. I have yet to see a single case where a usage report was audited by a third party.

Last week I wrote about a Chinese delivery platform handing out free camera helmets to riders. A line from that piece applies just as well here. Whether it’s labor or creative work, the danger point isn’t when compensation is low — it’s when the transaction stops looking like a transaction.

$656 might be a low figure. But at least it was a transaction with an amount written into a contract. Because the amount is written down, you can ask for a raise next time, compare it against other authors, or push the publisher on the revenue split. Negotiation stays possible.

Once you move to settlement based on reference counts, making those demands gets much harder. There’s no way for me to check how many times my book was referenced this month. When the settlement comes back at ₩0 (~$0), I can’t tell whether nobody actually referenced it or whether it just got dropped from the count.

So I don’t think what creators should be demanding right now is a higher per-unit rate. It’s these two things: a minimum guarantee and access to usage reports. Not a floor drawn by Spotify, but a floor that creators demand first.

Closing

There are three things that matter here.

  • AI training-data contracts have, until now, priced the mere fact of inclusion in the training set. HarperCollins’s $5,000 per book, Anthropic’s $3,000 per work in its settlement, and my own $656 — all of these set a single price per book or per work.
  • The usage-based model that Apple, Microsoft, and Perplexity are shifting toward prices content according to how many times it’s actually referenced. That structure fits news, which gets referenced fresh every day — but once it spreads to books and papers, the vast majority that are rarely referenced end up settling at close to $0.
  • We already have results from places that tried usage-based pricing before us. Spotify drew a floor line and cut songs with too few plays out of payouts entirely; the UK’s Public Lending Right set an annual cap that limits how much any one author can receive. Today’s AI contracts have no such correcting mechanisms built in.

If you happen to be negotiating an AI-related contract with a publisher or agency right now, I’d urge you to check just one sentence this week. Does the clause say compensation is calculated based on usage — and do you have the right to verify that usage yourself? Right now, a lot of contracts don’t spell out that right of verification at all.

💬 Have you ever faced a contract spelling out the terms under which your writing, photos, lectures, music, or code get used for AI training or services? I’m less curious about the dollar amount than about exactly what sentence was written into the contract. Was it a flat fee or usage-based? Was there any way to verify usage? Tell me in the comments — once I’ve gathered enough cases, I’ll use the next issue to sort them into types from the perspective of actual contract practice in Korea.


📨 If you know someone who writes, photographs, composes, or codes for a living, please pass this issue along to them.


Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

Past issues worth reading alongside this one


Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Grounding: A method where AI fetches up-to-date external material on the fly and uses it as the basis for its answer. Unlike training, which bakes knowledge into the model in advance, grounding pulls in fresh references every time a question comes in — which is exactly why “how many times it was used” can be counted.

  2. Long tail: The long stretch of items trailing behind a small number of popular hits — each one small on its own, but adding up to a large total. In the book market, this is the hundreds of thousands of titles that sit behind the bestsellers.

  3. Public Lending Right (PLR): A system that compensates authors a set amount each time their book is borrowed from a library. Because it prices usage rather than sales, it’s the oldest form of the usage-based pay we’re discussing today.