Issue #155

Gemini 3.6 Flash: Why Cost-Per-Task Beats Price-Per-Token

The real competition now is which model finishes the same job using fewer tokens, not which one charges less per token.

BusinessGemini 3.6 Flash: Why Cost-Per-Task Beats Price-Per-Token

The draft looks accurate overall. One check: the title says “17% Price Cut” but the Korean source doesn’t mention a price cut anywhere in the body—only pricing figures and 17% fewer tokens. Let me verify this is a heading match issue, not a fabricated fact.

Output Token Price Down 17%, Meeting a 17% Drop in Token Consumption

On July 21 (local time), Google announced Gemini 3.6 Flash. To be precise, it’s actually three models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. This falls outside my regular publishing schedule, but the timing felt too relevant to sit on, so I’m sending out this special edition.

Pricing is $1.50 per million input tokens and $7.50 per million output tokens. Most coverage has focused on the benchmark scores. But one other line in the announcement caught my eye.

“Consumes 17% fewer output tokens than the previous model.”

Token price alone doesn’t tell the whole story anymore — how few tokens it takes to get the same job done is becoming just as much a competitive benchmark.

What happened

Google’s rolled out three models this time: the flagship Gemini 3.6 Flash, the ultra-lightweight 3.5 Flash-Lite, and the security-focused 3.5 Flash Cyber.

Pricing table

3.6 Flash’s output price is $7.50 per million tokens, down roughly 17% from its predecessor 3.5 Flash’s $9.00. Input pricing stays frozen at $1.50. Performance climbed too: on DeepSWE, a benchmark for long-horizon coding tasks, it went from 37% to 49%; on MLE-Bench, a machine-learning engineering benchmark, from 49.7% to 63.9%; and on OSWorld-Verified, which tests direct manipulation of a computer screen, from 78.4% to 83.0%.

Google also emphasized efficiency gains. According to the Artificial Analysis index1 , output token consumption dropped 17% overall, and by as much as 65% on tasks like DeepSWE. In other words, the number of reasoning steps and tool calls themselves went down.

The companion 3.5 Flash-Lite is interesting in its own right. Priced at $0.30 for input and $2.50 for output, it churns out 350 tokens per second and scored 54.2% on the SWE-Bench Pro coding benchmark — beating the previous generation’s top model, 3 Flash, which scored 49.6%. A lower-tier model outperforming a flagship from the prior generation.

Near the end of the announcement, Google mentioned it has begun pretraining Gemini 4. Meanwhile, 3.5 Pro is currently in partner testing.

You have to look at unit price and per-task token consumption together

To understand this, you need to look at the price trajectory of the Flash line.

In June 2025, Gemini 2.5 Flash cost $0.30 for input and $2.50 for output. Then 3 Flash rose to $0.50 and $3.00, and this past May, 3.5 Flash climbed to $1.50 and $9.00. Input pricing quintupled in a single year.

This trend has been read as “the end of the cheap-AI era.” According to the research outfit Epoch AI, the cost of delivering the same performance falls 5-10x every year, yet vendors have started not passing those savings on to customers. Given reports that OpenAI’s 2024 revenue was roughly $3.7 billion against losses of roughly $5 billion, one interpretation is that a growth strategy built on absorbing heavy losses has bled into pricing policy. Still, loss figures alone can’t explain why any individual model’s price rose, or by how much.

googleBut this time, 3.6 Flash cut its unit price. It looks like a reversal of the price-hike trend above — but what’s more notable than the 17% cut itself is that token consumption for the same task dropped along with it.

Even in chat-style AI, cost is determined by multiplying unit price by usage. Agents repeat cycles of planning, calling tools, and reviewing results, so differences in usage can compound even further. Token costs are incurred even in reasoning steps that people never see in the final answer. That’s why you need to look not just at unit price, but at the input/output tokens and tool-call costs required to complete a single task.

If output price drops from $9 to $7.50 and output tokens also fall by 17%, output cost becomes roughly 69% of the previous figure — a roughly 31% reduction by the math. That’s an output-cost comparison assuming the same token savings as the announcement, for the same task. It doesn’t mean the total bill — including input tokens, caching, and tool calls — drops by 31%.

This announcement is a case where a price cut and a drop in token consumption happened together. You need to check whether both changes actually produce the same savings in your own workload.

Rivals’ flagship models have all converged on the same price band

Line up this weight class’s price sheets as of July, and here’s what you get. OpenAI’s GPT-5.6 Luna charges $1 for input and $6 for output; xAI’s Grok 4.5 charges $2 and $6; Anthropic’s Claude Sonnet 5 is running a temporary discount at $2 and $10. Sonnet 5 reverts to its list price of $3 and $15 starting in September. Just a year ago, frontier-level performance simply didn’t exist at this price point — now the flagship products of all three companies are clustered right here.

The benchmarks split the difference. By Google’s own comparison chart, Luna leads long-horizon coding work (DeepSWE) at 67%, 3.6 Flash leads computer-use tasks (OSWorld) at 83.0%, and Sonnet 5 posts the highest Elo2 score on GDPval-AA v2, a knowledge-work evaluation, at 1607. In other words, there’s no outright winner. Since performance varies by task, the real comparison has to happen among models that already clear your quality bar — on cost and efficiency.

So here’s how I see the picture playing out. Price sheets will keep converging. Differentiation will come from three places instead: first, the token efficiency we’ve discussed today; second, effective-discount mechanisms like caching and batch processing; and third, routing3 — swapping between cheaper and pricier models depending on the nature of the task. Structuring your usage to switch models based on task complexity is also a strategy worth buyers’ consideration.

Oswarld’s Lens

From my own experience designing pricing tables for B2B products while building GTM strategy, lowering the cost-per-task that buyers actually feel helps with retention. Google’s efficiency gains here could well appeal to customers running agents at scale. But we can’t simply assume that better token efficiency means fatter margins for Google. That requires knowing the vendor’s compute costs and overall usage volume.

If you’re a buyer, you can’t decide based on a per-model price comparison chart alone. Even at identical unit prices, actual costs can diverge several-fold depending on how many tokens a model consumes. In the end, you have to measure it against your own workload. Pick a handful of representative tasks you run frequently in-house, and measure the cost per task every time a new model comes out. It’s far more accurate than staring at a price sheet.

There’s one thing to watch out for, though. That 17% efficiency figure is an average across whatever benchmarks the vendor chose. This is a pattern I’ve seen repeatedly while working with data—these averages vary widely by workload. You could easily see efficiency improve dramatically on document summarization while token usage actually goes up on code generation—a complete reversal. The benchmark should be your own measured results, not the vendor’s average.

Closing

Google announced that it cut 3.6 Flash’s output pricing by roughly 17%, and that output token consumption on certain benchmarks dropped by 17% as well. Actual savings will vary by task. Buyers need to measure output quality alongside input, output, and tool costs together, for their own specific workloads.

This week, pick three or four representative tasks your team handles with AI and measure the per-task cost of the model you’re currently using — just once. It’ll make your next decision about switching models far easier.

If you’ve switched models in your own work, let me know in the comments whether your bill ever came in different from what you expected.


📨 If you know a colleague wrestling with AI costs, share this piece with them.


The draft matches the source well structurally and terminologically. No corrections needed.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • Google, “Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber”, Google Blog, 2026. Link ··· This is the primary source for today’s newsletter. Both the benchmark tables and pricing come straight from here.
  • Artificial Analysis, “Gemini 3.6 Flash: Intelligence, Performance & Price Analysis”, 2026. Link ··· This is where the 17% token-efficiency figure comes from — the most useful independent metric for comparing real-world costs across models.

Background

  • Ed Zitron, Exclusive: Leaked Documents Reveal OpenAI’s Brutal … Financials, 2026. ··· This confirms the base year behind the widely-reported 2024 revenue and loss figures.

  • XDA Developers, “Google’s Gemini 3.5 Flash costs 3x the model it replaced, and the era of cheap AI is ending”, 2026. Link ··· This piece lays out the pricing trajectory across the Flash lineup and the “end of the subsidy era” argument, which ties directly into Chapter 2 of today’s issue.

  • 9to5Google, “Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4”, 2026. Link ··· This article covers the announcement context and the teased news about Gemini 4 pretraining.

  • AI2Work, “GPT-5.6 vs Claude Sonnet 5: Inside the Frontier Price War”, 2026. Link ··· I referenced this for competitors’ pricing structures and information on time-limited discounts.


Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Footnotes

  1. Artificial Analysis Index: A benchmark service that independently measures and compares the performance, speed, and cost of major AI models. Because it’s based on third-party testing rather than vendors’ own claims, it’s widely cited across the industry.

  2. Elo score: A relative-ranking method borrowed from chess rating systems. It pits two models’ outputs against each other and raises the score of whichever “wins,” so it reflects relative standing rather than an absolute score.

  3. Routing: A method of automatically assigning an incoming task to the most appropriate AI model among several, based on the task’s difficulty. Easy tasks go to cheaper models and hard tasks go to pricier ones, lowering overall cost.