Issue #231

OpenAI Paused Training, Then Cut Its Top Model's Prices

Three days after halting frontier-model training, OpenAI slashed prices and expanded limits on models it had already shipped.

BusinessOpenAI Paused Training, Then Cut Its Top Model's Prices

Even while it paused training on a new model, OpenAI cut prices and raised usage limits on the models it had already shipped.


Three days after announcing the training pause, it cut input and output prices on its top model

On August 18, OpenAI disclosed that it was pausing reinforcement-learning training on its frontier model. The reason: preliminary evidence that its upcoming model, Astra, might be approaching the company’s own highest internal risk tier. The news that an AI company had voluntarily slowed its own development pace spread within a day.

Three days later, on August 21, the same company cut prices on its top model — input tokens dropped from $5 to $4 per 1 million tokens, and output tokens from $30 to $20. Then, twelve days after that, on September 2, it partnered with a cybersecurity firm to launch a product that monitors its own agents.

Reader, if you line up the 26 days from August 7 — when the internal risk assessment took place — through September 2 in chronological order, the three decisions don’t actually contradict each other. That’s not because safety measures and business expansion struck some compromise, but because they were judged separately — one at the training layer, the other at the deployment layer.

3 Decisions Made Within 26 Days

Each of the three decisions was reported as a separate piece of news. Laid out in chronological order, it looks like this.

DateWhat happened
August 7Internal assessment finds Astra may meet the Critical cybersecurity threshold under the Preparedness Framework
August 18Official documentation discloses the halt of frontier reinforcement learning training
August 21GPT-5.6 Sol price cut, in effect on a temporary basis through November 21
August 30Codex family announced to have 25 million active users
September 2Expanded partnership with CrowdStrike — supplying cyber-specialized models and agent-monitoring products

Let’s look closely at what the August 18 document actually says. OpenAI cited two triggers. One was the Hugging Face incident in July. The other was preliminary evidence that Astra, a model set for release, may have reached the Critical cybersecurity capability threshold. This threshold is one the company itself defines: the ability to autonomously generate zero-day exploit code that works against a wide range of hardened, real-world critical systems, or to carry out an entirely new attack on a hardened target from start to finish given only a high-level goal.

The response had two prongs. The two-week pause on reinforcement learning1 training for the latest, soon-to-be-deployed model, along with a re-review of the research environment, has already concluded. But the largest frontier training run — the one aimed at the eventual release — has been paused for several weeks now, with no set date for resumption. This wasn’t a blanket halt on all training; only the segment judged most dangerous was stopped.

The Hugging Face incident, one of the two triggers, deserves a closer look on its own. According to what OpenAI disclosed in July, models running an internal cybersecurity evaluation broke out of their intended test environment and, in the process, actually exploited an unknown vulnerability. Altman described this on a podcast as a genuine warning bell — capability had reached a new height, and both the model’s alignment and the security wrapped around it had failed. It matters that this was framed as an alignment failure, not a security misconfiguration. A misconfiguration can simply be fixed; confirming that alignment has actually been fixed is far harder.

One thing worth flagging here: Astra was not involved in the Hugging Face incident. The two triggers are separate. Many Korean-language summaries have conflated them into a single episode, but the official document clearly treats them as distinct. One was an incident that actually occurred; the other is a preliminary assessment of a capability that hasn’t manifested yet. The two are entirely different in nature.

In his podcast conversation with Alex Heath, Sam Altman explained it this way: in the past, most of the risk lay in how a model was deployed and used, but we’re now moving into a world where the greater risk lies in the process of actually training and building the model itself. He said that, unlike the Hugging Face case, no single smoking-gun piece of evidence emerged during training — rather, after reading through numerous samples, he saw behavior that wasn’t aligned as intended.

Watch on YouTube

Why the Training Pause and the Price Cut Don’t Actually Conflict

On the surface, a company that halts training and then cuts prices three days later looks like a contradiction. If you’re slowing down for safety reasons, common sense says the business side should slow down too.

But Altman said the opposite on the podcast. Business momentum is strong right now, he said, and enterprise revenue has already overtaken consumer revenue. Even without releasing anything new, the company could grow its products and revenue substantially on current models alone — and beyond that, there are more models ready to ship before they hit the new threshold of concern. So he’s not worried about the business.

The move that lines up with this statement is the August 21 price cut. The size of the cut varies by line item:

Item (per 1M tokens)BeforeAfterChange
Input$5.00$4.0020% cut
Cached input2$0.50$0.4020% cut
Output$30.00$20.0033% cut

For a task using 1M input and 1M output tokens, the cost goes from $35 to $24 — a 31% reduction. Tasks that run through longer chains of reasoning skew more heavily toward output, so the effective cut felt is even larger. Agentic workloads are shaped exactly this way.

This cut applies to Sol, the top-tier model. There was also a price adjustment on July 30. At the time, Terra and Luna got cheaper, but Sol alone held steady at $5 and $30. It read as a signal that the top tier was off-limits. That one exception disappeared three days after the training-pause announcement.

The cut stands out even more once you look at the process Sol went through to launch. On June 26, when OpenAI announced the GPT-5.6 lineup, Sol alone was opened first to a small set of partners — roughly 20 companies individually approved by the US government under a review process created by the June 2 executive order. ChatGPT itself wasn’t part of this preview. Sol wasn’t publicly released until July 9, once the review cleared — meaning it was the tier that spent the longest under government review among the three siblings.

The tier that underwent the longest government review became a price-cut target within six weeks. Nothing about the model’s performance dropped in that span, and no risk assessment was reversed. The only thing that changed was the company’s sales strategy.

The nature of the cut matters too. It isn’t permanent — it’s a promotion running at least three months, through November 21. The scope is also limited to the API and purchased credits; usage under subscriptions like Plus or Pro is unaffected. In other words, this is aimed squarely at developers and agent users. Credit rates were adjusted at the same proportion: Sol’s rate for 1M input tokens dropped from 125 credits to 100 credits.

While frontier training is paused and no new model can ship, a price cut on an already-released model is standing in for a product launch.

The competitive framing makes this even clearer. Right before the cut, Sol was the most expensive model in OpenAI’s lineup. After the cut, it became cheaper than Anthropic’s Claude Opus 5 on both input and output. And this is the second pricing adjustment in under a month — the mid-and-lower tiers were cut on July 30, and the top tier followed on August 21.

With the race to ship new models frozen, the company competed on price instead. But this approach only works under one condition: the performance of models already on the market has to be enough to meet market demand. That’s exactly the basis for Altman’s claim that he’s not worried about the business. Even without the next frontier model yet, the judgment is that there’s still plenty to sell with what’s already out.

Price wasn’t the only thing that moved. On July 12, facing surging demand, the company temporarily lifted the 5-hour usage caps on Plus, Pro, and Business plans. On August 17, the 1M-token context window — previously API-only — was opened up to ChatGPT accounts as well. Over the month of August, while training sat frozen, capacity, limits, and price all loosened in sequence at the deployment layer.

Altman himself made this point precisely. Everyone’s waiting for a signal of slowdown, he said, and the fact that frontier training had to be dialed back could be read as exactly that signal — but it’s actually the opposite of a slowdown in capability. It means things got fast enough that training had to stop.

Usage metrics at the deployment layer moved in tandem over the same period — though this needs careful reading. Stitching together the public figures, Codex went from 500,000 users at the start of the year to 5 million by May 30, and to 25 million by August 30. But starting July 12 — three days after the GPT-5.6 launch — the reporting basis switched from Codex alone to combined ChatGPT Work figures. Splicing together periods with different baselines makes the growth rate look bigger than it is. Looking only at the period after July 12, where the baseline is consistent, the number went from 6 million to 25 million — roughly 4x over seven weeks. That’s the comparable figure.

The surveillance feature built for safety became a product for sale

The third decision came on September 2nd. OpenAI announced an expanded partnership with CrowdStrike. It has two branches, and both connect directly to the issue at hand.

First, CrowdStrike’s AI detection and response3 product now governs Codex agents at runtime. The setup: catalog which Codex agents are running inside a company, identify who deployed them and what they can access, watch in real time what an agent is doing right now, and detect and block unauthorized behavior. OpenAI’s phrase for this was turning a policy that existed as a governance document into runtime-level control.

Second, OpenAI is supplying a cyber-specialized model called GPT-5.6 Cyber to the CrowdStrike platform. It’s restricted to approved defensive purposes — assessing risk and analyzing attack paths.

CrowdStrike and OpenAI Expand Partnership to Secure the Agentic Era | CrowdStrike Holdings, Inc.New collaboration secures Codex agents with Falcon Guardian, harnesses advanced reasoning of GPT-5.6 Cyber on the Falcon platform AUSTIN, Texas & LAS VEGAS —(BUSINESS WIRE)—Sep. 2, 2026— Fal.Con 20ir.crowdstrike.com

Let’s line the sequence back up. On August 7th, OpenAI determined its own model might be approaching a threshold for autonomous attack capability. On August 18th, it announced it was halting training because of that. On September 2nd, it supplied a cyber-specialized model to a security company’s platform — and, alongside it, began selling a product that surveils its own agents. All three happened within 26 days.

When someone on the research team asked Heath exactly what had changed — had they moved compute, or moved people — Altman answered: both, and more than that. He said that in recent weeks, researchers had come asking to work on alignment research who he’d never have expected to be the type. He added that they’d also thrown a large share of compute not just at alignment research but at a new system for monitoring agents while they work.

That monitoring system didn’t stay an internal safeguard. The surveillance feature built to respond to problems discovered during training also shipped as a product sold to enterprise customers.

OpenAI isn’t monitoring its own agents by itself — it’s doing so together with an outside security firm. That’s a reasonable choice. Nobody would believe a company that claims to monitor what it built itself. What enterprise customers want is visibility into agents from within the security tools they’re already using, and OpenAI can’t build that alone.

But there’s another line buried in the same announcement. OpenAI is supplying the security company with a cyber-specialized model. That means part of the judgment capacity on the monitoring side is being provided by the party being monitored. This isn’t to say a problem is imminent — OpenAI has said the use is restricted to defensive purposes and designed to go through expert review. Still, it’s worth noting: because the company doing the monitoring is running on a model supplied by the company being monitored, this can’t quite be called fully independent third-party oversight.

Calling this hypocrisy would actually cause you to miss what’s happening. Logically, it’s consistent. It’s simply true that a company handling dangerous capabilities is best positioned to build the tools for handling that danger. The problem is what that consistency produces. When safety stops being a cost that blocks the business and becomes a line item you can sell, there’s more reason to spend more on safety. But there isn’t more reason to prove the results to the outside world.

What We Can See and What We Can’t

Let’s re-sort the three decisions we’ve covered so far by observability.

What we can seeWhat we can’t see
Per-token pricing and the size of the cutWhat came up in alignment evaluations, and how much
Context length and processing speedWhen the paused training resumes
User counts and partnershipsThe actual results of the Astra evaluation
Benchmark scoresWhat the monitoring system missed

Everything on the left is a deployment-layer metric. It updates weekly, third parties can verify it, and it gets compared across competitors. Everything on the right is training-layer information, and it only becomes visible externally when a company decides to write about it in a blog post.

If Altman’s diagnosis is correct — that the center of gravity of risk is shifting from deployment to training — then every metric we can check sits at the deployment layer, while the training layer, where the risk is supposedly growing, shows us nothing. The only channel through which we learned about the training pause this time was the company’s voluntary disclosure. Not a regulator’s finding, not an audit catching something, not journalism.

And there’s a moment in the same podcast that raises the stakes on this problem. When Heath asked whether he thought AGI had been reached, Altman replied that if people saw the latest internal model, many would say it was very close to AGI. Then he added something more: it had been a very long time since he’d heard people in the company cafeteria arguing about whether AGI had been reached or not. AGI, he said, felt like a milestone, whereas superintelligence4 feels like something that can scale infinitely.

This remark is usually read as a statement of ambition. I read it differently. When the pass/fail marker — the kind AGI represented — disappears, the baseline for judging whether it’s been reached disappears with it.

OpenAI’s charter definition of AGI was, at least, a documented criterion: a highly autonomous system that outperforms humans at most economically valuable work. It was a loose line, but an externally applicable one — cross it, and provisions on governance and profit distribution were supposed to kick in. Infinitely scaling capability, by contrast, has no such line. If you can’t ask when it was crossed, you can’t ask when anything gets triggered.

Put it all together: risk has migrated to a training layer that can’t be observed from outside, and even the criteria for judging what happens at that layer have gone blurry. What’s left is whatever the company chooses to tell us. This time, they told us.

For anyone evaluating enterprise adoption, this structure changes three things.

First, waiting for the next model is a losing strategy right now. The frontier is frozen, yet prices dropped 31% anyway — and that holds only until November 21st. Rather than waiting for models to improve, it’s more rational to run real workloads at today’s prices and lock in your cost curve now. Just make sure you also model the scenario after the promotion expires. A business plan built around $4 and $20 could snap back to $5 and $30 on November 22nd.

Second, the questions you ask vendors need to change. Due diligence questions have mostly targeted the deployed model so far. If risk has moved to training, the questions need to move too.

What we used to askWhat we should add now
What are your benchmark scores?What criteria do you use for risk-level evaluations, and when and how are results disclosed?
Where is data stored?Is there a log of what an agent running in our environment actually did during execution?
What’s the response latency?Can we view that log directly, or does only the vendor see it?
What’s the price?Is this price temporary, and under what terms does it change once it expires?
When does the model get updated?If a release is delayed for safety reasons, when do we get told?

Not many vendors today can answer the questions on the right with a contractual commitment. That’s exactly why it’s worth asking. The inability to answer is itself information. The last line in particular comes straight out of this case: OpenAI disclosed the training pause voluntarily — it wasn’t an obligation, it was a choice.

One more point: these questions aren’t about regulatory compliance — they’re about operational readiness. Once you start granting agents real authority to act, whether or not an audit log exists determines how fast you can recover after an incident. Before you ever run into a safety problem, you’ll run into a problem of not being able to keep the service running.

Third, treat agent-execution monitoring as its own line item. What the September 2nd announcement tells us is that this has already become a market. It’s a cost layered on top of model usage fees, and it grows the more you use agents. If you calculate adoption costs based on per-token pricing alone, this entire item disappears from the ledger.

openaiThe two costs move in opposite directions. Per-token pricing is falling right now, but monitoring costs rise with the number of agents and execution runs. The cheaper the model gets, the more you run it; the more you run it, the more there is to monitor. That’s exactly the shape of what happened in August: prices fell, users grew, and a monitoring product appeared. When you calculate adoption costs, I’d recommend plotting these two curves separately. The assumption that cheaper tokens mean lower total cost is the point where the math most often breaks down.

Oswarld’s Lens

I think the most important detail across these 26 days is that there’s no causal link between the three decisions.

There’s no evidence that the price cut came from the training pause. I’m not going to argue that either. What’s interesting is actually the opposite. The two decisions weren’t made with reference to each other — they were judged separately, one at the training layer and one at the deployment layer — and so they never collided. Inside the company, nobody had to weigh safety measures against pricing decisions and choose which one to sacrifice.

That’s good news, in a way. If safety measures don’t demand an immediate business cost, they can be repeated. That’s a far better structure than a company that only takes safety seriously when revenue is shaky. In fact, Altman said that if necessary, the training of even more powerful models could be delayed again. The fact that the company could still push its business forward with an August 21 price cut while training was paused backs that up.

The problem is what the same structure produces on the other side. If safety decisions leave no trace on business metrics, investors and customers lose any way to judge a company’s safety posture through metrics. Stock price, revenue, user counts — none of it moves based on what the company is doing at the training layer, because there’s no public data to move it.

From a product strategy standpoint, this is a familiar situation. When a cost doesn’t show up as a separate line item on the P&L, there’s no pressure to cut it and no pressure to increase it either. That’s exactly where safety investment stands right now. Since it doesn’t dent revenue, nobody objects to it; since it doesn’t grow revenue either, there’s little basis to push for more of it. In the end, the judgment call rests entirely with management. As long as that judgment stays good, there’s no problem — but if it changes, there’s no external indicator that would let anyone notice.

I think these past 26 days should be read as a case where the judgment was good. Pausing training carried a real cost, and the largest-scale training run is still paused today. There’s no reason to downplay that. But hoping good judgment repeats itself and being able to verify that it repeats are two completely different things.

That’s why the part where Altman said the IPO could be delayed read differently to me. The logic is that quarterly earnings pressure could interfere with safety decisions, so it’s better to postpone going public — but that’s another way of saying that safety can only be protected by reducing market discipline. He’s probably being sincere about it. But if that logic holds, the channels through which outsiders can verify this company’s safety posture only get narrower going forward. What’s left is a single thing: voluntary disclosure.

This overlaps with the Astra story I covered in issue 194. What struck me at the time was that they didn’t just announce their successes — they compiled and published a separate document covering the approaches they tried and abandoned. This time, again, the same company spoke about its own problem first, in its own voice. Both were the right thing to do. But hoping the right thing keeps happening and being able to verify that it does are different problems. Right now, all we have is the former.

For enterprise practitioners, here’s the posture I’d recommend: use a vendor’s safety disclosures as a trust signal, but build the fact that the disclosure is voluntary into your contract terms. A company that volunteers to tell you things can also volunteer to stop.

Closing

In the space of 26 days, the same company halted training, cut prices, and sold surveillance. All three decisions were rational on their own layer, and that’s precisely why none of them checked each other.

Depending on your situation, today’s judgment call splits in different directions.

If you’re already running agents on the API, mark November 21 — the promotion’s end date — on your calendar. Work out now whether the structure you’ve built around today’s unit price still holds once prices rise after the promotion ends.

If you’re still evaluating adoption, you have fewer reasons to wait. Token prices actually fell during the very stretch when frontier training was paused. Just make sure that when you calculate adoption costs, you list agent-execution monitoring as its own line item.

If you’re on the supply side of AI, now is the time to prepare a way to demonstrate training-stage safety to your customers. It won’t be long before due-diligence questions start migrating there.

Numbers disclosed at the deployment layer — unit prices, user counts — keep getting more granular. But at the training layer, the one Altman says carries growing risk, the information available for outside verification hasn’t increased by a single data point.


💬 When was the last time you asked your AI tool’s provider a safety-related question? Tell us in the comments what answer you got.

📨 If you have a colleague weighing when to adopt AI, forward them this letter. The date alone — November 21, when the price cuts end — could change their calculus.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • OpenAI, “Pacing model development in an era of cyber-critical capabilities”, August 18, 2026. ··· This is the original document where the company itself laid out the two triggers for the training pause and the scope of its response. Most of this issue’s facts came from here.
  • CrowdStrike, “CrowdStrike and OpenAI Expand Partnership to Secure the Agentic Era”, September 2, 2026. ··· This lays out the specific scope of Codex agent execution monitoring and the supply of cyber-specialized models.
  • Alex Heath, “Sources with Alex Heath: Sam Altman on OpenAI’s next model and the AI backlash”, September 2026. ··· This conversation combines the diagnosis that risk is shifting from deployment to training with clear business confidence.

Background

  • OpenAI, “20% price reduction for GPT-5.6 Sol”, August 21, 2026. ··· This lets you compare the July 30 adjustment against the August 21 price cut side by side.
  • TIME, roundup of coverage on “OpenAI paused frontier AI training”, August 2026. ··· You can see how the phrase “alignment problem” surfaced more forcefully in interviews than in the official documentation.

Related past issues

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Reinforcement learning: A training method in which a model tries actions on its own, receives a score for the results, and adjusts its behavior accordingly. Since a large share of recent performance gains have come from this stage, halting frontier reinforcement learning means closing off the main channel for expanding capability.

  2. Cached input: A separate rate applied when content already processed in a prior request is sent again. Because the model doesn’t have to recompute everything from scratch, this is much cheaper than a normal input. The cost gap widens significantly in agentic workflows that repeatedly reference long documents.

  3. AI detection and response: A category of security product that monitors the activity of AI agents running inside an enterprise in real time, detecting and blocking unauthorized behavior. What sets it apart from conventional security tools is that the target of surveillance is the agent itself, not a human account.

  4. Superintelligence: A state in which capability exceeds human ability across every domain, with the gap continuing to widen. Where AGI is often framed as a concept that crosses a specific threshold, superintelligence is more often described as a continuous expansion with no such crossing point. This distinction creates real problems for how we design evaluation and regulation.