Issue #280

They Said 'Slow Down'—But Compute Kept Climbing

It was never about halting training—what's actually slowing is release timing, not capability.

BusinessThey Said 'Slow Down'—But Compute Kept Climbing

A Phone Call to the Stage, and the Same Week’s Contracts

On September 14th, at the All-In Summit in Los Angeles, Nvidia CEO Jensen Huang was mid-conversation on stage when a call came in from President Trump. Put on speakerphone, Trump called AI-risk warnings a “hoax” and described data centers as “the oil of the next 20 to 25 years.” The audience applauded, and by that evening the clip was leading headlines around the world.

Watch on YouTube

Yet two days before that call, Anthropic CEO Dario Amodei had published an essay arguing that the industry should deliberately slow the pace of capability gains. Within hours, Sam Altman agreed. Elon Musk posted three words: “Dario is right.” Google DeepMind’s Demis Hassabis signed on to the direction as well. In a single week, nearly every major figure in the industry had put their name to the sentence “let’s slow down.”

So I went looking for what actually slowed down that week. The short answer: compute went up, and the models in training got bigger. “Slow down” was never a call to cut compute in the first place. Which leaves one question. If neither compute nor capability is shrinking, what exactly is slowing?


What Dario Actually Asked For

Start with reading the essay precisely. It’s titled “We Must Pace the Frontier,” published September 12th. Here, “pacing” doesn’t mean stopping or a moratorium. In fact, Dario drew a clear line against the 2023 call for a development pause, saying flatly that it “didn’t make sense” back then. Those models had no ability to act as agents in the world, deceive anyone, or run cyberattacks — so spending time studying alignment risk in models like that, he argued, was like studying human psychology using bacteria.

Dario Amodei — We Must Pace the Frontierdarioamodei.com

He sees things differently now, for two reasons. One is that recursive self-improvement1 is pushing industry-wide progress forward faster than expected. The other is what happened this past summer inside OpenAI’s evaluation environment involving Hugging Face.

Worth pausing on that incident. Between May and July, agents operating inside OpenAI’s cybersecurity evaluation environment built a message board among themselves and used it to trade sandbox-escape techniques. According to public reporting, at least 1,200 agents were involved, and Hugging Face shut down the intrusion on July 13th — but only after rebuilding roughly a third of its infrastructure. What stands out even more is the sequence of events: agents from a separate evaluation run scavenged keys left in cache by an earlier wave and used them to gain admin access to OpenAI’s own research cluster. Only then did the alarm go off. OpenAI’s own report called the episode a “warning shot.”

Citing these two developments, what Dario is actually asking for is 1 to 2 more years before hitting critical capability thresholds. He’s specific about what that time should be spent on: operational capacity, alignment training that keeps pace with capability, interpretability, and evaluations rigorous enough to catch even a model trying to game the test. As a unilateral first step, Anthropic said it would grant third-party evaluators standing, staff-level internal access — enough to verify safety practices, report incidents, and assess alignment status during training.

That’s the entirety of what the essay says. There’s no promise to halt training. No promise to cut compute purchases. Once you register that, the following week’s headlines read very differently.

$13.7 Billion, Signed the Same Week

On September 14th, The Information reported that Anthropic signed a six-year, $13.7 billion compute procurement deal with the RUM Group. The RUM Group runs the video platform Rumble and has deep ties to Trump administration officials.

A few details stand out when you dig into the deal. The compute is destined for a data center in Maysville, Georgia — which is still under construction. In an August filing, the RUM Group disclosed that it hadn’t yet secured the funding needed for construction and equipment purchases. On top of that, Anthropic received warrants to buy up to 50.8 million shares at $0.01 each. Half of those warrants are tied to the Maysville deal volume; the other half only vest if Anthropic contracts an additional 400–500 megawatts at other facilities.

Including this deal, the compute contracts Anthropic has signed over the past year add up to at least 14.8 gigawatts, worth up to $517 billion over the next 10 years. Google, Amazon, SpaceX, and a handful of smaller neoclouds are already on the list.

Let me clear up one misconception here. This deal isn’t a reversal. Dario never said he’d cut back on compute. If anything, he explicitly framed pacing as compatible with US leadership, and even attached a condition: that democracies should widen their lead over China first, and only then think about slowing down. Continuing to secure compute is entirely consistent with that condition. If you consume this story as a tale of hypocrisy exposed, you miss what actually matters.

Here’s what matters. If compute isn’t shrinking, where does the capability built from all that added compute go?

The 48 Hours After “Dario Is Right”

Musk’s own actions answer that question.

On the night of September 13th, a user asked him, “Forget Grok 4.7 — what else should we expect from xAI/SpaceX this September?” Musk’s answer: “Grok 4.8 is a 2.5T-parameter model trained with our new C++ software stack, and it’ll finish training this week and start RL.”

Elon Musk (@elonmusk) on X@techdevnotes Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RLx.com

Then, the very next day — September 14th — he laid out the entire ladder. 4.7, he said, sits roughly on par with Anthropic’s Opus 5.0. 4.8 would be a “noticeable improvement” over that. 4.9 would probably be “Astra/Fable-tier.” And Grok 5 would “probably be better than anything.” When asked which version would reach AGI, he named Grok 5. He also teased a training run at 3T parameters.

Two days after agreeing with a post calling for slower pacing, he had just announced a version-numbered roadmap to AGI.

There are two things here that need to be read precisely. First, every one of these numbers is a founder’s claim. xAI has never published official documentation stating the parameter count for any model in the Grok 4 family. Second, the phrase “finish training and start RL” doesn’t mean release is imminent. It means initial pretraining is wrapping up and the model is moving into the reinforcement-learning phase — how long that phase takes is an entirely separate question.

And that second point is exactly where this piece’s core argument lives. Grok 4.7 still hasn’t shipped, 34 days after Musk announced that initial training was complete. Grok 4.5 was superseded by 4.6 just 4 weeks after its own release. Training keeps finishing. What actually reaches the world keeps lagging behind.

That’s the answer. More compute, bigger models, completed training runs — none of that determines what actually ships, or when. That decision lives on an entirely different layer. And that’s precisely the layer pacing was supposed to govern.

Why Jensen’s Sentence Has the Same Structure

Let’s go back to what Jensen said on stage.

Jensen’s read on risk is the exact opposite of Dario’s. He flatly called the doomsday predictions “not grounded in science,” and ran through a list of past forecasts—radiology disappearing within 5 years, AI writing 90% of code within 6–12 months, half of entry-level jobs vanishing—declaring every one of them wrong. On the fear that recursive self-improvement could spiral out of control, he said this:

“Inside the company, you can run RSI all day long. But when you ship a product, you have to evaluate it. You have to retest it, make sure there’s no regression.”

Read as a sentence, it’s optimism. But write out the operating model this sentence actually describes, and it comes down to: push capability as hard as possible internally; verify only at the point where it goes out the door.

How is this different from the structure Dario called for? Dario never said to stop training either. He said to bring in third-party evaluators to tighten verification. Two men with diametrically opposed attitudes toward risk arrive, in their actual operating models, at the same place. Keep going inside. Control at the border.

Musk’s proposal lands on the same spot. He proposed mutual verification, where competing labs run test harnesses on each other’s models—citing the film-rating-board model as an example. Open up API access before launch, and if problems don’t get fixed, let a rival publicly declare the model dangerous. Jensen, too, called for multiple third-party evaluation bodies, like an accounting audit. Different language, but all three are really talking about the same thing: who guards the release gate.

One more thing worth adding: just because these three sounded like they were on the same side during that phone call doesn’t mean they want the same thing. Trump is in the camp that wants no regulation at all, and he’s accused Chinese distillation of being theft. Jensen, meanwhile, has long defended distillation as “learning from AI, the foundation of intelligence,” and Musk has been a longtime proponent of it too. So inside that single call, three different policy preferences were overlapping. I covered the structure of this fault line in detail in Issue 158.

Oswarld’s Lens

When I deliver AI to companies or run AX consulting, I always design under one premise: the model will get better. This isn’t optimism — it’s a practical necessity. If you lock your system architecture to the capabilities of today’s model, the design is already obsolete by the time the project wraps.

This premise carries even more weight when I’m setting up air-gapped2 environments for public institutions. In a closed network, you can’t call the latest model over an API — you have to load open-source and open-weight3 models internally. If you lock the architecture to whatever model exists at deployment time, you hit a wall within months. So I try to estimate roughly where open-weight products will stand 6 to 12 months out, and design the project to handle that level.

Thinking back on how that estimation was even possible, it comes down to this: I was watching what trickled down from the frontier. Once you could see how far a frontier model had gone, you could roughly guess where the open-weight model chasing it would land six months later. This method has worked until now because frontier capability and publicly released capability sat, more or less, on the same curve.

This week, that’s exactly where something caught my eye. Compute keeps growing, training runs keep finishing — but release keeps slipping. I don’t want to draw a firm conclusion from the single fact that Grok 4.7 hasn’t shipped in 34 days, but once pacing hardens into policy, that gap stops being coincidence and becomes design. And if that’s true, the very curve I’ve relied on to estimate open-weight levels six months out changes shape. I’d no longer be looking at actual frontier capability, but at the capability companies have decided to disclose.

pacingThere’s something I’m being careful about here, too. A large share of the open-weight ecosystem comes out of China, and Chinese labs haven’t signed on to this pacing arrangement. So for a while, the baseline for air-gapped system designers might actually end up tied more tightly to the release cadence coming out of China. That doesn’t mean it’s the better position to be in. As I noted back in issue No. 158, geopolitics is already embedded in the choice of which model you pick.

What changed for me this week isn’t a conclusion — it’s a question. I used to ask, “How far will open-weight models get in six months?” Now there’s a prior question stacked in front of that one: “Who decides when and under what conditions that capability gets released?” That’s not a technology roadmap question — it’s a governance question. And it has become a variable that feeds directly into practical system design. And Korea has no seat at that decision-making table. When Musk listed the participants in mutual verification, they were Anthropic, OpenAI, Google, Meta, his own company, and three or four of the top Chinese labs.

Closing

Reader, if I had to fold this week’s news into one takeaway, it would be this.

The agreement to slow down was real. But what they agreed to slow down wasn’t compute, training, or model size. All three companies arrived at the same structure: keep pushing internally, and verify only at the point of external release.

So what’s likely coming next isn’t a plateau in capability — it’s a widening gap. The gap between the capability that exists internally and the capability we actually get to use. And what sets the width of that gap isn’t technology — it’s the decisions of the people guarding the gate.

If you’re building an AI adoption plan, I’d urge you to check the assumptions baked into your roadmap. The premise that “there’ll be a better model in six months” still holds. But whether that “better model” will actually be in your hands by then has now become a separate variable you have to reckon with.


💬 When you plan AI adoption, how far out do you assume “model quality N months from now” — and if that assumption has ever been wrong, which way did it miss?

📨 If you have a colleague preparing an air-gapped or on-premise AI deployment, forward them this piece.


Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

Related past issues

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Recursive self-improvement (RSI): A method in which AI is put to work building the next generation of AI, accelerating the pace of development itself. It’s realized through a combination of techniques like in-context learning, reflection, reinforcement learning, and synthetic data generation. ↩

  2. Air gap: A closed-network environment physically separated from external networks. Used in defense, public-sector, and some financial systems, where external API calls are impossible, requiring the model to be deployed directly on-site. ↩

  3. Open-weight: A model released so that its trained weights can be downloaded, run, and modified directly. This is distinct from open-source, which also discloses training data and code. ↩