Aleph Logo

ALEPH

AboutContact

AI Budgets Disappear in 4 Months — The Truth Behind the Surge in AI Operating Costs, Seen Through the Uber Case

💸 AI Cost Analysis
AI operating costs
Claude Code
Thinking Tax
Uber AI
AI ROI

It is true that AI triples productivity. The problem is that no one warns us in advance that costs will skyrocket even faster. Uber's CTO admitted that the company exhausted its 2026 AI budget in just four months . This was the result of introducing Claude Code to 5,000 engineers. This is not just Uber's story. Let's analyze in detail the structural trap where costs grow exponentially the more AI is used.

AI productivity-to-cost explosion structure
Productivity rises linearly, while costs rise exponentially. This difference is not visible in the early stages of AI adoption, but at some point, the budget runs out first.

What Actually Happened at Uber — The Full Story of Burning Out a 4-Month Budget

The core of the problem regarding the skyrocketing AI operating costs lies in the speed at which usage is exceeding expectations. According to a report by The Information (April 2026), Uber CTO Praveen Neppalli Naga admitted that the 2026 AI budget was depleted much faster than originally planned. The main culprit identified as the cause was Claude Code .

Approximately 95% of Uber's engineers have adopted Claude Code, and 70% of their total code is generated by AI. In terms of productivity metrics alone, this is a successful case of AI adoption. However, the surge in usage resulting from this productivity improvement led to the exhaustion of an annual R&D budget of $340 million in just four months.

Claude Code adoption rate of engineers
95%
Targeting approximately 5,000 total Uber software engineers (The Information)

AI-generated code ratio
70%
The proportion of Uber's total code written by AI (AIMagazine)

AI budget early exhaustion period
4 months
Period when the annual AI budget was exhausted — Uber CTO remarks (The Information)

AI Project ROI Shortage Rate
80~95%
Rate of AI projects with under-ROI or cancellation — Industry estimates (MIT, IBM)

2Thinking Tax — A structure where money goes out the more AI thinks

Thinking Tax refers to a billing structure where costs accumulate at each step as AI performs the reasoning process. Unlike typical API calls, token costs are accumulated separately at each reasoning step where the latest AI models analyze complex problems step-by-step.

In agentic loops that involve the repetitive process of writing, reviewing, and debugging code, like Claude Code, this "Thinking Tax" accumulates even more steeply. The estimated cost for Claude Code is $500 to $2,000 per engineer per month; simply applying this to 5,000 engineers results in a scale of $2.5 million to $10 million per month. This is the mathematical reason why annual AI budgets are exhausted early. The current structure involves a simultaneous drop in unit costs and a surge in usage; even if unit costs decrease, total costs continue to rise if usage increases even faster.

Cost Types General API calls Agentic + Reasoning Loop
Billing units Input/Output Token Input / Output + Inference Step Tokens
Single task cost $0.001~0.01 $0.05~1.00 (varies depending on task complexity)
Repeating loop effect linear increase Exponential increase (rapid increase depending on the number of loops)
Cost Predictability height Low — Large variation depending on task complexity

Thinking Tax Cost Accumulation Structure Diagram
A Thinking Tax structure where costs accumulate as the AI goes through inference stages. Unlike general API calls, costs in the Agentic loop accumulate exponentially at each inference step.

Why 3 Is a carbon copy of the Big Data era — History repeats itself

The problem of skyrocketing AI costs is structurally identical to the failure pattern of the big data boom 10 years ago. At that time, too, companies built Hadoop clusters, constructed data lakes, and invested tens of billions under the conviction that “data is the future.” The results are exactly as recorded in reports by MIT and IBM— 80–95% of AI and big data projects failed to demonstrate ROI or were discontinued. These figures are industry estimates, and it is appropriate to interpret them as a direction rather than precise values.

The same mistakes are being repeated even in the era of AI agents. As the simplistic framework of "replacing employees with AI" is applied without governance and cost planning, companies are facing a reality where productivity metrics rise but budgets are exhausted prematurely. The key is not simple replacement, but cost planning, orchestration , and an ROI measurement framework .

division Big Data Era (2010~2018) AI Agent Era (2024–Present)
Initial expectations Data solves all problems AI replaces employees and reduces costs
Actual result Data Lake Construction → Minimal Utilization Increased productivity → Soaring costs
cause of failure Lack of ROI design, lack of governance Thinking Tax, usage prediction failure
ROI underperformance ratio Over 85% (IBM estimate — industry estimate) 80~95% (MIT · S&P Global — Industry estimate)

4 A reality approaching both companies and individuals

The skyrocketing operating costs of AI operate differently regardless of company size. Larger enterprises have higher usage, while smaller enterprises lack the buffers to absorb the cost shock.

From a corporate perspective, exceeding AI budgets leads to project cancellations and a failure to recoup ROI. For individual developers, the tangible experience of triple productivity with a single Claude Code is followed by API bills rising faster than expected. NVIDIA executives stated, “AI computing costs have started to exceed employee salaries.” By the second half of 2026 , “AI agents that operate stably for more than a month” are expected to become the new benchmark for technological competitiveness.

5 Practical Strategies for Controlling Costs — 3 Prescriptions

The key to solving the problem of AI operating costs is not reducing usage, but designing the cost structure . The following three strategies are currently the most effective approaches.

1

Introduce an Orchestration Layer — Design the Cost Flow

Set a Task Budget (maximum cost limit per task) for all tasks where agentic loops occur. By utilizing the A2A (Agent-to-Agent) protocol to monitor the tokens consumed by each AI agent and the inference steps, you can identify where the Thinking Tax spikes and block it in advance. Cost predictability is the first condition of AI governance .

2

Multi-vendor Strategy — Reduce Dependence on a Single Model

It is not necessary to use the highest-performance model for every task. There are reports that a hierarchical model strategy , which uses low-cost models (such as Haiku and Gemini Flash) for repetitive and simple tasks and selectively employs high-performance models only for tasks requiring complex inference, can reduce costs by 30–60%. This also simultaneously prevents vendor lock-in.

3

Redesigning ROI Metrics — From Productivity to Cost per Outcome

Measuring that “code output increased after the introduction of AI” is incomplete. To calculate actual ROI , Cost per Outcome , Completion Rate , and Time to Value must be tracked simultaneously. The moment cost metrics are overlooked while focusing solely on productivity indicators, a situation like Uber's repeats itself.

6Frequently Asked Questions

question answer
What exactly is Thinking Tax? This is the additional token cost incurred at each step when an AI model reasones on a complex problem. Models that perform step-by-step reasoning, such as Claude Extended Thinking, consume tens to hundreds of times more tokens than simple API calls.
Does the Uber case apply to domestic companies as well? Structurally, they are identical. Only the scale differs; even domestic companies that have adopted tools like Claude Code and GitHub Copilot on a team basis are exposed to the risk of exceeding their budget due to failed usage forecasts. Even with just 10 engineers, monthly AI costs can range from $5,000 to $20,000.
Won't AI costs go down in the future? Model unit costs are trending downward. However, as the use of Agentic Loops and Reasoning functions increases, the cost per task actually rises. Due to the structure where unit costs drop and usage surges simultaneously, total costs are highly likely to continue increasing.
Where should I start with AI cost governance? The first step is to establish a real-time monitoring system for current AI usage and costs. Optimization itself is impossible without identifying which teams or tasks are concentrating costs.

3-Step Strategy for Optimizing AI Operating Costs
AI cost optimization begins with designing the cost structure, not reducing usage. Building an orchestration layer, a multi-vendor strategy, and measuring Cost per Outcome are the three key prescriptions.

Conclusion — The winner of the AI era is not where it is used the most.

The core message conveyed by the skyrocketing operating costs is simple: AI is not a matter of adopting productivity tools, but of designing the cost structure . As demonstrated by the case of Uber, if a structure is left unchecked where costs rise even faster than productivity metrics, the budget will eventually hit rock bottom.

To avoid repeating the failures of the big data era, you must ask yourself two questions right now: Are we tracking our team's AI usage in real time? And is the cost of adopting AI greater than the cost savings? If you cannot answer these questions, exhausting your budget is only a matter of time.

Paradoxically, this cost crisis creates new opportunities. Engineers and companies capable of designing cost-effective AI architectures become the key assets in the next AI competition. The game has begun where the winner is not the one using the most expensive models, but the one operating them most efficiently.

⚠️ Investment Precautions
All figures in this article are for informational purposes only and do not constitute investment advice. Uber-related figures are based on reports from sources such as The Information, Yahoo Finance, and AIMagazine, as well as industry estimates, and may differ from official financial disclosures. The underperformance ratio (80–95%) and personal expense estimates are industry estimates; please interpret them as general trends rather than exact figures. You bear all investment decisions and responsibilities.

📌 Was this analysis helpful?

In the next post, we plan to cover “Winners in the Thinking Tax Era — Companies and Strategies Creating a Competitive Advantage with AI Cost Efficiency.”

We continuously track the intersection of AI cost structures and investment strategies. Receive notifications so you don't miss out.

Subscribe and get notifications →

How this content was produced

Aleph's research AI agent assisted with collecting and analyzing public data, creating charts and visuals, and structuring the draft. Davar personally reviewed and edited the sources, figures, reasoning, and final conclusions.

This content is for informational purposes only and is not personalized investment advice or an individual stock recommendation. Read the full disclaimer

Aleph Logo
Founder · Author · Editor

Davar

Davar builds and operates Aleph's research AI agent and writes and reviews analysis on macroeconomic developments and AI industry trends.

공유하기

© Aleph. All rights reserved.