Issue #225

Inference Margins Hit 80%. Why Hasn't AI Pricing Moved?

I separate what AI firms report as margin from what cloud providers actually charge for inference.

BusinessInference Margins Hit 80%. Why Hasn't AI Pricing Moved?

Cloud Takes $40 Off the Top—So Why Is the Margin 80%?

When an AI company earns $100, $35 to $40 of it flows straight into another company’s bank account. That other company is AWS, Azure, or Google Cloud. It’s the compute cost of running inference1. The three cloud giants keep $10 to $20 of that as operating profit—an operating margin of roughly 35 to 45%.

These are the numbers from Barclays’ AI industry unit economics report, published on August 28, 2026. I checked the industry summary table myself—everything in it is a Barclays estimate. Figures for 2026 and beyond are forecasts, not realized results.

The same report includes another detail. It estimates that AI labs’ margin on paid inference will rise from a low-double-digit percentage in 2025 to over 50-65% in 2026, with adjusted gross margin improving by 30 to 50 percentage points within a single year. For direct API sales, the estimated inference margin tops 80%.

Here’s the puzzle. If $40 out of every $100 goes to compute costs, how can the margin be 80%? The two numbers use different denominators. The $40 is cloud cost calculated against the AI lab’s total revenue, while the 80% is the ratio of (API revenue minus the direct compute cost of that specific request) to that same API revenue. Below, I’ll walk through why the price buyers pay hasn’t budged even as margins have climbed this much.

Why the Two Numbers Seem to Clash

Let’s start by separating the denominators behind these two figures. Here’s Barclays’ industry summary, reproduced as-is.

Unit: $B202420252026E2027E2028E
AI lab revenue726137376690
Inference cost
(% of revenue)
4 (57%)17 (66%)58 (42%)157 (42%)292 (42%)
Training cost
(% of revenue)
7 (96%)19 (70%)66 (48%)132 (35%)210 (30%)
Hyperscaler AI revenue1136124289502
As % of AI lab revenue153%136%90%77%73%

These are Barclays Research estimates; figures from 2026 onward are forecasts.

Look at the 2026 column: inference cost sits at 42% of revenue, training cost at 48%. Add those together and you get $124B — which is exactly the hyperscalers’ AI revenue. That’s 90% of AI lab revenue.

The “$35-40 out of every $100” figure only isolates the inference share. Training is excluded.

“80% inference margin” narrows the frame even further. It takes API revenue earned from running an already-trained model and subtracts only the direct compute cost of serving those requests. Training costs, research headcount, data costs — none of that enters the picture.

So the first number is the company’s overall P&L, and the second is the contribution margin of a single production line. Both are correct, but they don’t belong on the same axis. For reference, ICONIQ’s tally of actual gross margins for AI products comes in at 41% in 2024, 45% in 2025, and roughly 52% in 2026. That’s improving, but it’s nowhere near 80%.

The 2024 column makes this structure even clearer. That year, combined inference and training compute costs equaled 153% of revenue — compute costs alone exceeded revenue. That was the era when hyperscalers were earning more than the AI labs themselves.

This table is also easy to misread. A common interpretation goes something like “90% of hyperscaler AI revenue comes from AI labs” — but that gets the direction backward. The denominator here is AI lab revenue. It means that for every $1 an AI lab earns, a hyperscaler earns 90 cents — not that 90% of hyperscaler revenue comes from AI labs. This table tells you nothing about the latter. Whenever you read a table like this, the first thing to check is what’s sitting in the denominator.

Subscriptions Are the Least Profitable Product

Barclays broke down inference margins by product line: direct API over 80%, indirect API somewhere in between, and subscription products around 70%. Flat-rate subscription products like Claude Code or Codex have the lowest inference margin of the three lines.

You’d think bundling a flat-rate subscription would yield better margins, but the result is the opposite.

The reason is simple: AI labs are deliberately subsidizing the token costs of subscription users. These are the users they need to retain. API customers have already built products on top of the model, making it hard for them to switch — but subscription users can just cancel next month.

The recent trend of subscription usage limits resetting frequently can be read the same way. Part of it is due to improved model efficiency, but it also looks like pressure to prevent churn is working in tandem. This is Barclays’ observation, not an established fact.

To sum up: AI pricing is being set less by cost and more by which customers need to be retained. The subsidy is concentrated not on the line with the highest cost burden, but on subscription users who are most likely to churn.

Different business mixes make the books look different

Barclays built two hypothetical frontier labs to compare. Neither is a real company—both are models.

Lab A gets 70% of revenue from API and 30% from subscriptions. Lab B is the reverse: 80% subscription, 20% API. Same industry, similar products—yet Lab A’s adjusted gross margin comes out around 55%, versus about 38% for Lab B. That’s a 17pp gap.

Half of that gap comes from the margin mix you just saw; the other half comes from accounting. Lab A books indirect API revenue on a gross basis2, while Lab B books it net—or doesn’t recognize the strategic partner’s operated share as revenue at all. Barclays compared this to the relationship between Uber and Lyft: the same underlying business, but different recognition standards make the reported numbers look different.

Flip to the cloud side, and the picture inverts. For every $100 of Lab A’s revenue, cloud revenue comes to $35, operating profit $11.8, a 34% margin. Lab B, with revenue-sharing layered in, shows $41 in cloud revenue, $19.1 in profit, and a 47% margin.

Barclays attaches a caveat here: what revenue-sharing inflates is the headline margin, not the actual money kept per token—that’s the same in both cases. And they expect that revenue-share arrangement to disappear after 2028.

One more point worth flagging. Agentic subscription products generate additional value on the cloud side. Because a request doesn’t end in a single call but keeps a task’s state alive through continuous execution, it pulls in shared resources like databases and storage along the way. That means more consumption rides along with every dollar of revenue. In some segments, there are even contracts where the cloud provider and the AI lab split this consumption take.

The same logic applies on the adoption side, too. The more agents you deploy, the more you’re paying not just in model call fees but in the infrastructure consumption trailing behind them. If a quote only lists the per-call model price, the infrastructure consumption cost that follows is missing from it.

Once AI labs start disclosing GAAP financial statements, investors will have to confront these revenue-recognition differences directly. We’ve already covered a moment when metrics became hard to take at face value—that time, the problem was how ARR was calculated. This time, it’s the denominator of margin.

But Cost Doesn’t Move in Just One Direction

Everything so far has assumed that “cost is falling.” Per-token inference cost really is falling. Fewer tokens are needed to finish the same task, techniques like quantization and speculative decoding have been layered on, and a new generation of compute has arrived.

But the price of buying that compute is moving the opposite way. This trend is getting stronger as high-end components like M8/M9-grade CCL3 become the mainstream.

CategoryPrice increaseNote
Memory+45% q-qQ2 2026
CCL+12~22%M8/M9 high-end
MLB+15%For AI accelerators
Package substrate+15%For memory
MLCC+15~35%High-capacity, for AI servers
AI servers+15% or more

This is data Samsung Securities compiled from press reports. The memory price increase is lower than the 58~63% for commodity DRAM that TrendForce projected for Q2, which suggests it’s closer to a blended realized figure across product lines.

Two different costs are moving in opposite directions. The cost of producing a single token is falling, while the cost of buying the machines that produce tokens is rising. And this increase hasn’t fully hit AI labs’ income statements yet, because servers are bought once and then flow through as depreciation over several years.

So it’s more accurate to see today’s 80% not as leftover slack, but as the buffer that will have to absorb the rising procurement costs still working their way into future earnings.

Why It Matters

Three practical implications follow from this structure.

First, don’t build price cuts into your budget as a given. When organizations draft AI adoption plans, many assume “token prices will keep falling anyway” and project three years forward on that basis. But over the past year, while margins improved by 30–50 percentage points, the price cuts buyers actually received fell well short of that. The cost declines didn’t flow through to customers — they stayed inside the company. In some segments, nominal API prices even went up.

Barclays’ forecast points the same direction. Inference cost as a share of revenue drops sharply from 66% in 2025 to 42% in 2026 — but then stays flat at 42% through 2027 and 2028. The bank is essentially saying the big cost declines are already behind us. What improves from here is the training-cost share, not the per-unit inference price.

Second, a vendor’s business mix tells you where you have room to negotiate. Vendors with a high share of API revenue have more margin cushion, but also a stronger incentive to defend price. Vendors with a high share of subscription revenue run thinner margins but are more sensitive to churn. With subscription-heavy vendors, threatening to cancel works better as a negotiating lever than offering volume commitments. The same ask can get you a different answer depending on which type of vendor you’re facing.

Third, the cloud’s share of the pie is projected to keep shrinking relative to AI lab revenue. The ratio of hyperscaler AI revenue to AI lab revenue falls from 153% in 2024 to 90% in 2026, 77% in 2027, and 73% in 2028. The training-spend share follows the same path, dropping from 96% in 2024 to 35% in 2027 and 30% in 2028. Even so, the absolute dollar figures keep growing: hyperscaler AI revenue rises from $124 billion in 2026 to $502 billion in 2028. What’s shrinking is the ratio relative to AI lab revenue — the dollar amount itself keeps climbing.

make moneyStill, the direction is clear. AI lab revenue is growing faster than hyperscaler AI revenue, and starting in 2028, committed in-house infrastructure projects come online. Any cloud infrastructure contract or investment decision needs to account for this timeline.

Industry-wide revenue is projected to grow from $7 billion in 2024 to $137 billion in 2026 and $690 billion in 2028. These are all forecasts — if this scale doesn’t materialize, the ratios above collapse along with it.

Closing

I want to leave you with three things.

First, when you look at AI margin numbers, check the denominator first. An 80% inference margin and a 52% gross margin are both correct figures. So is 90%. All you need to check is what’s being divided by what.

Second, today’s prices are set less by cost calculations than by the need to hold onto customers. The subsidies are concentrated on the subscription side, where churn risk is highest.

Third, costs don’t move in just one direction. While the per-token cost keeps falling, server and component prices are rising — and that rise will flow into future earnings through depreciation.

Have you actually seen the unit price of the AI tools you use drop over the past year? Or did only your usage limits go up while the bill stayed the same? I’m curious which one it was for you.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Glossary

Footnotes

  1. Inference. The process of actually running an already-trained model to generate answers. Unlike training, which builds the model, inference keeps happening as long as the service is live.

  2. Gross vs. net revenue recognition. In brokerage-type transactions, this is the difference between booking the full transaction value as revenue (gross) versus booking only the fee (net). The same business can look like it has a different revenue scale and margin depending on which method is used.

  3. CCL. Copper clad laminate, the base material for printed circuit boards. MLB refers to multi-layer boards, and MLCC refers to the small multi-layer ceramic capacitors used in circuits.