Issue #115

Why Uber Capped AI Coding Tools at $1,500 a Month

Uber caps AI coding tool spend at $1,500 per employee monthly, spotlighting how firms must weigh cost against value as usage grows.

BusinessWhy Uber Capped AI Coding Tools at $1,500 a Month

Does spending more tokens actually mean better results

Back in March, Jensen Huang made a case on the All-In podcast — recorded during GTC — that engineers should be given a generous budget for AI usage. His suggestion: an engineer earning $500,000 a year should be able to spend something like $250,000 annually on tokens. The idea was to push people to use the tools aggressively enough that they’d get more done. Coverage of the remarks

On June 2, it came out that Uber has been running a $1,500-per-employee, per-AI-coding-tool monthly usage cap. A company spokesperson confirmed the details to Bloomberg. The limit was apparently rolled out over the past few months. There’s no confirmed causal link suggesting Uber followed Jensen Huang’s advice and then walked it back — that’s not established anywhere. But taken together, the two stories show something real: companies are actively wrestling with how to manage costs even as they push employees to use AI more. Bloomberg’s report

Jensen Huang’s Budget Proposal Isn’t a Performance Formula

What Jensen Huang is really arguing is that if a high-salary engineer can produce bigger results, there’s no need to skimp on AI costs. He isn’t presenting experimental proof that spending and productivity always move in lockstep.

Divide the annual token1 budget of $250,000 by 12 months, and you get roughly $20,833 a month. This is a proposal to allow for a large budget — not a cost estimate saying every developer must spend exactly that much. The amount actually needed depends on the task at hand, how much time AI saves, and the value of the output.

We also need to keep in mind that Nvidia supplies the compute infrastructure behind AI. As AI usage grows, demand for Nvidia’s products can grow with it. But not all of a client company’s token spending flows back to Nvidia as revenue — model providers, cloud services, and software costs all take their share.

Companies adopting AI, like Uber or Walmart, have to weigh what they spend against what they get. If rising token usage doesn’t translate into better service or lower costs, it becomes hard to keep justifying the investment.

Financial firms sell financial services, retailers sell goods, automakers sell cars — AI is simply a tool that helps them do that job. The real yardstick for any adopting company should be how much value customers receive and how much operating costs change as a result.

Infrastructure suppliers and the companies that adopt their tools don’t always have opposing interests. When better tools lead to better performance, both sides can benefit. The problem arises when usage itself gets treated as performance. In a culture of so-called token-maxing2 — pushing token consumption as high as possible — even unnecessary calls can end up being encouraged.

Uber Scaled Adoption and Spending Controls at the Same Time

Uber is a case study in how quickly an engineering organization can scale up its use of AI coding tools.

According to reports, Uber rolled out Claude Code3 to an engineering organization of roughly 5,000 people in December 2025, alongside tools like Cursor4. The company reportedly also ran internal leaderboards tracking usage. This rollout predates Jensen Huang’s March 2026 remarks.

Coverage of Uber cites figures showing 95% monthly adoption of AI tools among developers, with roughly 70% of commits receiving AI assistance. These numbers demonstrate that the tools are widely used. But the share of AI-assisted commits shouldn’t be read as meaning 70% of all code was written without human involvement. Forbes’ rundown of the rollout

In April, it emerged that Uber had burned through its annual budget for AI coding tools like Claude Code faster than expected. This shouldn’t be stretched to mean the company’s entire AI investment budget — including its existing dispatch and pricing prediction models — had been wiped out. Nor can you multiply one developer’s monthly spending range by headcount to arrive at a definitive total cost.

The caps Bloomberg confirmed are independent per agentic coding5 tool. Spending $1,500 on Claude Code doesn’t shrink the Cursor limit in tandem. In some cases, the caps can reportedly be exceeded with approval. So $1,500 shouldn’t be treated as one employee’s entire AI budget, nor compared directly to other companies’ broader annual budget proposals.

There’s still the matter of connecting usage to customer value. Uber President and COO Andrew Macdonald said on the Rapid Response podcast that it’s hard to explain the relationship between improving AI usage metrics and delivering more useful features to consumers.

A record of developers producing more code is a different thing from customers actually having an easier time using the service. The latter can only be confirmed by separately measuring real-world usage and quality.

This comment shouldn’t be read as an admission that AI has no value at all. It means that adoption and usage metrics alone are insufficient to explain the benefit that actually reaches customers.

We need to separate the volume of code from the quality of service

This concern isn’t limited to Uber.

Bloomberg reported that Walmart has also started limiting usage of its internal AI tools. Companies appear to be encouraging adoption while simultaneously setting budget caps. The mere existence of a cap doesn’t mean the rollout failed, but it can be read as a signal that the phase where expanding usage alone was enough is coming to an end. Bloomberg’s report on corporate AI costs

Menlo Ventures estimated the 2025 market for generative AI coding tools at roughly $4 billion. That’s about 55% of the $7.3 billion spent on department-level work AI. It doesn’t mean coding tools account for 55% of companies’ overall AI budgets. This is a report that combines a survey of US corporate decision-makers with market estimates, and that scope needs to be kept in mind. Menlo Ventures 2025 report

Given how much money is flowing into coding tools, we need to include the cost of reviewing and testing the code after it’s produced.

CloudBees’ 2026 report surveyed more than 200 enterprise technology leaders. The self-assessed readiness score for AI-generated code averaged 83.6 out of 100. That doesn’t mean 83% of respondents said they were ready. Separately, 81% of respondents reported an increase in operational service issues tied to AI-generated code. This isn’t a comprehensive measurement of actual failure rates, but it does show confidence and operational difficulty appearing side by side.

Weekly PR6 counts, lines of code, and commit counts are metrics that show volume of work. Now that AI can generate large amounts of code, these numbers alone have become harder to use for comparing developer performance. Review time, revision counts, post-deployment errors, and customer feature-usage rates all need to be examined together. This doesn’t mean existing metrics have become entirely useless—it means we shouldn’t assign too much meaning to any single metric.

Oswarld’s Lens

What bothers me most about Jensen Huang’s budget proposal is that it leaves out how to verify results. Even if you allow more spending, there still needs to be a standard for evaluating what that spending produces.

I’ve seen this pattern before in building tech strategy: teams adopt a tool quickly, and only after costs balloon do they try to measure the return on investment. The same thing happened with CRM and cloud computing. With AI coding tools, if you scale up usage first, you end up discovering — belatedly — which tasks drove up costs and what actually improved.

Writing code is just one part of software development. Even if coding gets faster, if design review or testing still takes as long as before, the overall release velocity may not speed up proportionally. It’s the same logic as the Theory of Constraints7: you first have to identify the bottleneck that’s limiting overall performance. Beyond token usage, you need to look at how long it actually takes to complete a task, and the quality of the output.

Uber didn’t stop using the tool — it set a cost cap. What matters going forward is whether the same cost produces better results. Even when raising that cap, I think you need evidence not just that more tokens were used, but that total cost — including review time — actually went down for specific tasks.

Closing

An AI budget should be designed to let teams use tools freely while still being able to verify results. Uber’s case shows that as usage scales quickly, cost management and impact measurement have to move together. Beyond counting tokens and lines of code, you need to check how those outputs actually reached customers.

How is your organization measuring the impact of AI tools? And are you perhaps experiencing your own version of “token-maxing”? Let’s talk about it in the comments.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Footnotes

  1. Token: The basic unit AI uses to process text. It doesn’t map one-to-one to letters or words, and its length varies depending on the language and the model’s tokenizer. AI service costs are typically billed based on this token consumption.

  2. Tokenmaxxing: A neologism for a corporate culture that pushes AI token usage as high as possible. Under the assumption that “more usage means more productivity,” some companies have even run usage leaderboards.

  3. Claude Code: An agentic coding tool made by Anthropic. It goes beyond simple code suggestions, autonomously writing, editing, and running code from the terminal.

  4. Cursor: A development tool that integrates AI directly into a code editor. Built on VS Code, it lets AI analyze and modify code at the whole-file level.

  5. Agentic Coding: A coding approach in which AI doesn’t just suggest a single line of code, but autonomously plans and executes multi-step tasks — creating files, running tests, and even fixing errors on its own.

  6. PR (Pull Request): The process by which a developer submits written code for team review. It’s the peer-review step code goes through before being merged into the main project.

  7. Theory of Constraints: The theory that a system’s overall performance is determined by its slowest bottleneck. Even if one segment is improved, if another segment remains the bottleneck, the overall gain in performance is limited.