Aleph Logo

ALEPH

Can AI Beat Inflation? — The Answer Lies in the “Inference Unit Price”

When Inference Unit Costs Fall, It Is Margins, Not Sales, That Waver — Inflation Series Part 4

📉 AI Disinflation
Inference unit price
Chinese open weight
OpenRouter
Token economy
AI investment

I left a question at the end of my last post: Is AI really beating inflation? When Fed Chair Wash brought up AI productivity before interest rates at the first FOMC meeting, what he was looking at was labor productivity. The idea was that if the same person produces more in the same amount of time, price pressure decreases. However, when I actually tried to find traces of this in the data, they were hard to spot. Productivity is slow to measure and slow to spread throughout the entire economy. Instead, another number was moving much faster and more clearly visible: the price per token.

OpenRouter US and China AI Model Token Market Share Trends

In OpenRouter's weekly token usage, the Chinese model rapidly overtook the US model after early 2026. This observation suggests the possibility that cost-sensitive inference demand is shifting. (Source: Financial Times, based on OpenRouter data)

This article does not aim to discuss the U.S.-China struggle for AI hegemony . I will set aside the question of who builds the smarter models for a moment. Instead, I will focus on just one thing: Is the unit cost of inference for performing the same tasks truly falling, and where is that decline occurring most rapidly? If Wash focused on labor productivity, what is moving more clearly these days is this unit cost of inference.

There is not just one way for 1AI to lower prices

There are actually two branches to the argument that AI suppresses inflation. One is the labor productivity aspect mentioned by Wash. If AI assists humans in their work to produce more output at the same cost, unit costs decrease, and price pressure weakens. This is the textbook path. The other is less discussed but is more direct: the drop in the unit cost of inference. If the same level of inference can be purchased at a lower cost, the cost of all tasks utilizing AI decreases along with it.

Although these two appear similar, their operating paths differ. The productivity path requires passing through people and organizations to yield results, which is why it is slow. While it certainly works in certain job functions, it takes time for it to spread throughout the entire economy. As pointed out in Part 1 , "Why Is the Market So Obsessed with AI?" , there are quite a few studies suggesting that the productivity effects of AI are smaller than expected or difficult to measure. On the other hand, the inference cost path does not require passing through people or organizations. If the cost per token changes, the cost price changes starting from that very day.

Here, we need to address a common misconception. News of Micron's strong earnings is sometimes interpreted as evidence that AI is lowering prices. That is a different story. Strong memory sales signal strong demand for AI, not that AI is driving down costs. In fact, in the short term, investment is concentrated in sectors like memory, power, data centers , and cooling, which tends to push costs up. For AI to act as a disinflationary factor, today's data center investment must translate into cheaper inference tomorrow. To determine if this is happening, one must look at token prices, not chip performance.

Cost-sensitive demand that moved first in 2OpenRouter

Where should one look to analyze token prices? One of the fastest-observed signals is OpenRouter. It is an intermediary platform where developers can select from hundreds of models using a single API, and the flow of traffic to these models reveals, at the very least, what cost-sensitive users are actually choosing. And that trend has rapidly reversed over the past year.

Throughout 2025, the US model drove OpenRouter's token usage, accounting for nearly 70% of the top models. However, between February 9 and 15, 2026, the Chinese model surpassed the US model for the first time. During that week, the Chinese model processed 4.12 trillion tokens, while the US model processed 2.94 trillion tokens. This was not a temporary reversal. By June, the gap had actually widened, with weekly usage based on a sample of the top nine models compiled by the Financial Times using OpenRouter data reaching approximately 18 trillion tokens for China and 5.5 trillion tokens for the US. This represents a ratio exceeding 3 to 1.

Why did this happen? It is not because the Chinese model suddenly overwhelmed the US model in terms of performance. Rather, the proportion of cost-sensitive tasks increased, and those tasks began to be routed to cheaper models. The share of coding and agent workloads on the platform rose from 11% in early 2025 to over half by mid-2026. Agents repeat a loop of planning, writing code, executing, and fixing dozens to hundreds of times. This is the same structure we saw in the article on token economics , "Claude Pricing Separation — The Investment Map Drawn by the Agent Token Explosion." In such tasks, the unit price of a token is no longer just a cost item, but the entire budget. If the same work can be done at a much lower unit cost, many developers prioritize the cost difference over minor performance differences.

Let us also clarify a point of caution. Commonly cited figures, such as "61% usage of Chinese models," refer to the proportion of token usage accounted for by Chinese models within the total token usage of the top 10 most used models on OpenRouter during a specific week. It does not represent the total number of over 400 models viewed across the entire platform, nor does it reflect revenue share. Nevertheless, the direction is clear. In a cost-sensitive inference market, demand is rapidly shifting toward low-cost inference.

To clarify, the term "open model" in this article does not refer only to open source in the strict sense. It is a collective expression that includes open weights, which allow for uploading publicly available weights to one's own infrastructure or routing them at low cost.

OpenRouter Weekly Token (Parent Model)
18 trillion vs 5.5 trillion
Based on a sample of the top 9 models compiled by the Financial Times using OpenRouter data — Chinese models approximately 18 trillion tokens, US models approximately 5.5 trillion tokens, reversing US lead from January (June 2026)

Performance gap between top US and Chinese models
2.7%
Based on Arena, the gap narrowed from 17.5–31.6%p in 2023 — top US models slightly outperform top Chinese models (Stanford HAI 2026 AI Index, March 2026)

Open Model Token Weight (OpenRouter)
34% → 65%
Rising between January and June 2026 — Cost-cutting trends shift toward open weight (Citi Note, cited by Reuters, June 2026)

Goldman Sachs Token Growth Outlook
24 times
Global Token Consumption in 2030 Compared to 2026 — Driven by Agent Proliferation (Goldman Sachs, citing FT)

3When the performance gap narrows, price determines it

At this point, a counterargument arises. I understand it is cheap, but if the quality is poor, won't it ultimately become unusable? Until last year, that was true. However, the situation has changed recently. The Stanford HAI 2026 AI Index explains that, based on major benchmarks, the performance gap between top-tier U.S. and Chinese models has narrowed to 2.7% by March 2026. This is a significant increase from 2023, when the gap stood between 17.5 and 31.6 percentage points across major benchmarks. Top performance remains close to the realm of U.S. frontier models. However, the sufficiently high quality required for high-volume iterative tasks is no longer a moat exclusive to the United States.

When quality becomes similar and prices drop significantly, cost-conscious parties are the first to make a move. This is already happening. According to a Reuters report on June 29, executives such as Microsoft's Nadella, Palo Alto Networks' Arora, and Coinbase's Armstrong have begun stating that a significant portion of enterprise demand can be handled by smaller and cheaper models. For a while, using a large amount of tokens was regarded as proof of productivity, but now that the bills are becoming a burden, the calculus is shifting. Based on Citi notes, the proportion of open model-based tokens on OpenRouter rose from 34% in January to 65% in June.

What is interesting is the contradiction within the cost structure. While the unit price per token continues to fall, the actual cost of completing a task is rising. This is because the number of steps an agent goes through to complete a single task increases, and the amount of data and input handled becomes longer. The Uber case discussed earlier illustrates this. As employees flocked to AI coding tools, the company exhausted its 2026 AI budget in just four months. Therefore, rather than being complacent about lower unit prices, there is growing pressure to switch to models that handle the same work at a lower cost. This is the actual reality of the disinflation currently occurring in the inference market.

Trend of narrowing performance gap between US and China AI models
The performance gap between top-tier U.S. and Chinese models narrowed from double-digit percentage points in 2023 to 2.7% by March 2026. As quality became similar, cost-sensitive demand began to shift toward the cheaper option. (Source: Stanford HAI 2026 AI Index)

4 The problem with the US frontier model is margin, not revenue.

So, are U.S. AI companies collapsing? That would be an exaggeration. To be precise, it is not that revenue is declining, but rather that inference margins and pricing power are being suppressed. The U.S. frontier model still leads in peak performance, enterprise trust, security, and cloud integration. Revenue per token is also higher than the Chinese model. The problem lies in the fact that it is becoming increasingly difficult to justify that premium.

It is easier to understand if you divide work into two types. There are tasks that absolutely require top performance, and those that do not. Massively repetitive work, such as drafting, data organization, and coding assistance, falls into the latter category. If a significant portion of this work flows to the cheaper open-weight model, the frontier model's ability to command a high price narrows. If the price remains the same, volume is lost; to retain volume, the price must be lowered. Either way, it reaches the margin.

In fact, signs of this are visible. According to a Reuters report, OpenAI is considering price reductions, including token usage fees. Anthropic is also not immune to the same price competition pressure. As both companies must simultaneously demonstrate high growth potential and profitability narratives, a price war places a burden on both revenue growth rates and margin stories. Palo Alto's Arora wrote in X: “If you want to win enterprise, you should be forward pricing tokens.” This means that instead of charging high prices now, they should apply prices that will drop in a few years. Essentially, the idea that one must lower prices first to secure corporate clients has begun to be openly discussed within the industry.

5However, this picture is easy to exaggerate

So far, we have only looked at the side with the strongest trends. However, this data is easy to misinterpret. Therefore, there are three clues that need to be examined together.

First, OpenRouter does not represent the entire AI market. It does not capture corporate traffic directly using OpenAI , Anthropic, Azure , or AWS, nor does it include usage from the government, finance, and security sectors. Due to the nature of the platform, it is skewed toward individual developers and cost-sensitive users. Therefore, stating that "the Chinese model has dominated the entire AI market" is inaccurate. The correct description is that rapid penetration is being observed in the cost-sensitive inference market.

Second, token usage does not equate to revenue. Even if the Chinese model processes more tokens, its revenue may be smaller than the US model if the unit price is low. Therefore, from an investment perspective, it is more accurate to say that "margins and pricing power are under pressure" rather than "US AI companies' revenue is about to collapse." Determining the revenue structure based solely on a usage chart leads one astray.

Third, just because something is cheap does not mean it can be used. As Reuters pointed out, while China's open-weight model is widely used by startups, it often blocks entry for large corporations due to security concerns. However, a distinction is necessary here as well. When directly calling APIs from Chinese providers, one must separately verify data processing locations, terms and conditions, jurisdiction, and security review issues. On the other hand, the risk structure changes when using publicly available weights on proprietary infrastructure. Inference based on low cost is not the same issue as inference based on compliance with regulations. As long as this shadow remains, it is difficult to view the entire market as being skewed in one direction solely by price competition.

6Frequently Asked Questions

question answer
Is it okay to use a Chinese model for work? The usage must be differentiated. There are significant cost advantages for personal work or non-sensitive coding that can be disclosed without issue. However, the risk varies depending on the call method. When directly calling APIs from Chinese providers, data processing locations, terms and conditions, jurisdiction, and security reviews must be verified separately. The structure changes when utilizing publicly available weights on proprietary infrastructure. Cheap inference and regulatory-compliant inference are not the same issue.
Should I sell US AI company stocks now? This article is not an investment recommendation. However, the key point of this trend is not a collapse in revenue, but rather pressure on margins and pricing power. It is better to differentiate between companies with a large proportion of business areas where frontier performance is essential and those heavily reliant on high-volume repetitive tasks. The first step is to dissect and examine the revenue structure rather than the company as a whole.
If inference unit costs drop, won't the demand for infrastructure like Nvidia also decrease? So far, it has moved in the opposite direction. During the two years when the unit price of tokens fell, consumption actually exploded. It is a structure where people use more as the price gets cheaper. This is an AI version of the Jevons Paradox seen in the previous article on the token economy. It is difficult to view a drop in unit price as directly leading to a decrease in infrastructure demand.
What does this mean from the perspective of Korean investors? The primary point is that lowering the cost of using AI reduces the financial burden on companies that utilize it as a tool. At the same time, price competition in the global inference market influences the pricing of domestic AI services. Either way, this trend appears to be shifting the standard from "smarter models" to "who can best utilize cheaper inference."
When can we verify if AI really beats prices? This is not the kind of thing that can be confirmed all at once. It is better to consistently observe a few indicators rather than making predictions. These include the token weight of open models, the unit price gap between US and Chinese models, and the speed at which companies actually switch routing to cheaper models. If these three move in the same direction, it can be seen as a signal that the power of AI to reduce costs is growing.

AI Disinflation Observation Indicators
The power of AI to suppress prices cannot be confirmed all at once. It is better to wait and see if these three indicators—the token weight of open models, the unit price gap between US and Chinese models, and the speed of routing transitions by companies—move in the same direction.

Conclusion — AI disinflation is first seen in token unit prices

When Wash brought up AI at the FOMC, he had labor productivity in mind. The idea is that if people work more efficiently, prices will fall. While this is true, it takes a long time for the effects to be reflected in prices, and it is still difficult to verify with data. In the meantime, what has been moving faster is the unit price of inference. Competition to supply inference of the same quality for cheaper tokens is changing the direction of actual demand.

So, to answer the question raised by this series—is AI really beating inflation?—we must change our perspective. Not on how smart the US Frontier model is, but on how quickly the Chinese and OpenWeight models are driving down the price per token. The conditions for AI to become a cost-reducing technology are surprisingly simple: models that handle the same tasks more cheaply must increase, and companies simply need to switch their routing to those models.

It is not yet time to draw definitive conclusions. The data comes from a single source, OpenRouter; usage does not equate to revenue, and the barrier of security remains. However, the indicators to observe have become clear: whether the token weighting of the open model continues to rise, how the unit price gap between the US and China widens or narrows, and whether companies are genuinely shifting routing to the cheaper option. If we consistently monitor these three factors instead of making predictions, the outline of the answer regarding whether AI is outpacing inflation begins to emerge.

⚠️ Investment Precautions
All figures and analyses in this article are for informational purposes only and do not constitute investment advice. OpenRouter data reflects a portion of the cost-sensitive inference market and does not represent the entire AI market. Cited figures are based on press releases and reports, and items involving estimates or interpretations are subject to further change. You bear full responsibility for your own investment decisions regarding the companies and stocks mentioned, and we recommend consulting a professional financial advisor before making any significant decisions.

📌 Was this analysis helpful?

In the next installment, we plan to cover “Is AI Productivity Really Lowering Prices? — Inflation Series Part 5.” This time, we will look for traces of labor productivity in the data, another path we set aside.

We continue to track the intersection of macroeconomics and AI. Subscribe to notifications so you don't miss out.

Subscribe and get notifications →

How this content was produced

Aleph's research AI agent assisted with collecting and analyzing public data, creating charts and visuals, and structuring the draft. Davar personally reviewed and edited the sources, figures, reasoning, and final conclusions.

This content is for informational purposes only and is not personalized investment advice or an individual stock recommendation. Read the full disclaimer

Aleph Logo
Founder · Author · Editor

Davar

Davar builds and operates Aleph's research AI agent and writes and reviews analysis on macroeconomic developments and AI industry trends.

공유하기

© Aleph. All rights reserved.