Claude plan
Agent AI
NVIDIA
Token economy
AI Infrastructure Investment
2026.05.14
Anthropic has completely separated agent usage and chat limits for Claude Max subscribers. While ostensibly a pricing overhaul, this small change encapsulates the entire next five years of the AI infrastructure industry. Agents consume up to 100 times more tokens than chats on an always-on basis. Furthermore, Jensen Huang has already declared this at GTC 2026 and the All-In podcast: “If a $500,000 engineer doesn’t use $250,000 worth of tokens, I will be very worried.” When the billing structure changes, the infrastructure landscape changes as well.

1I don't know what happened — the rate plan split in two.
Anthropic is completely separating Agent SDK-only credits from chat limits for Claude Max subscribers (effective June 15). The key point is not just a simple pricing adjustment. The key point is that “Conversational AI” and “Execution AI” have been separated into entirely different products.
| flan | Monthly subscription fee | Agent SDK separate credit | After running out of credits |
|---|---|---|---|
| Pro | $20 | $20 (Monthly reset, cannot be carried over) | API Pay-as-you-go conversion |
| Max 5x | $100 | $100 (Monthly reset, cannot be carried over) | API Pay-as-you-go conversion |
| Max 20x | $200 | $200 (Monthly reset, cannot be carried over) | API Pay-as-you-go conversion |
The numbers explain why this separation was inevitable. As agent tools like OpenClaw began leveraging Claude subscription limits, it became commonplace to consume $1,000 to $5,000 worth of computing power on a $200 per month Max plan . VentureBeat called this "compute arbitrage," and for Anthropic , it was a financially unsustainable structure.
Anthropic first blocked subscription authentication for third-party agent tools on April 4 and adopted a method of restoring them to a separate credit pool starting June 15. Reactions from the developer community are mixed. Some view it as a “effective price hike,” while others interpret it as the “beginning of transparent agent billing.” Regardless of how it is interpreted, the conclusion is the same: it is proof that agents have moved beyond chatting and entered the core engine of enterprise operations .
How many times as many tokens does the 2 agent eat compared to chat?
The “Decoding the Agentic Economy” report, released by Goldman Sachs on May 8, 2026, summarized the token consumption structure of agents in specific figures.
There is a structural reason why agents consume tokens in this manner. While chatting involves a single round trip of question to answer , agents repeat a loop of goal setting, planning, tool invocation, result verification, modification, and re-execution dozens to hundreds of times. According to Goldman Sachs estimates, a single email management agent consumes approximately 114,000 tokens per day. This is a natural consequence given that the system operates 24 hours a day without interruption.

3 Jensen Huang Already Predicted — Tokens Are Profit
At the GTC keynote in March 2026, Jensen Huang forecasted a demand for AI infrastructure worth $1 trillion by 2027. Immediately after GTC, he directly answered the questions investors were most curious about on the All-In podcast. Those who have read the commentary, which Aleph has summarized verbatim from the original text, will remember this remark.
“At the end of the year, I will ask the $500,000 engineer how much he spent on tokens. If that engineer hasn't spent at least $ 250,000 worth of tokens, I will be very worried.”
— Jensen Huang, All-In Podcast (2026.03.19) · Aleph Original Text Commentary →
Jensen Huang's logic is clear: token consumption is an indicator of a company's AI utilization, and that utilization generates revenue. The separation of Claude's pricing plans is proof that this equation has been implemented on a real business model. As agents began running 24/7 as the enterprise's operational engine, token consumption exploded, necessitating a complete overhaul of the pricing structure. As highlighted in the in-depth analysis of the GTC 2026 keynote , this demand ultimately becomes the engine that drives the entire AI infrastructure.
The logic of the energy market, where consumption increases as prices fall, applies here directly. Although the unit price of tokens has dropped to one-tenth of its level over the past two years, consumption has increased by more than 100 times. As the cost curve goes down, the infrastructure demand curve goes up.
Who Laughs When NUM3 Tokens Explode — Beneficiary Map
The benefit structure of the explosion in token consumption is divided into three layers. The magnitude and nature of the benefits vary depending on which layer one is in.
| Layer | Major companies | Benefit logic | personality |
|---|---|---|---|
| Layer 1 Chips and Infrastructure | NVIDIA, TSMC , SK Hynix , Samsung Electronics | Blackwell Ultra offers 50x higher performance and 35x lower cost per token compared to Hopper for agent workloads ( NVIDIA official). Lower costs lead to an explosion in usage (Jevons' Paradox). HBM is directly linked to agent long-context processing. | direct benefits |
| Layer 2 cloud | AWS, Azure , GCP | Surge in demand for agent inference computing directly translates to AI cloud revenue. AI inference market projected to reach $117.8 billion in 2026 and $312.7 billion in 2030 (Fortune Business Insights). | Intermediate beneficiary |
| Layer 3 AI platform | Anthropic, OpenAI | ARR increases directly as agent usage grows. Anthropic officially announced that its ARR will surpass $30 billion on April 7, 2026 (based on Anthropic's announcement). Goldman Sachs forecasts 24x growth in token consumption by 2030. | Structural benefits |
3 Real Signs That NUM4 Plan Changes Tell Us
This change to the Claude pricing plan points to three things simultaneously.
The agent has already gone beyond chatting.
The fact that a $200 plan consumes $5,000 worth of computing power has become commonplace means that agents have started being deployed in real production workloads. The usage gap has widened enough to require separate pricing. Agents are not an upgraded version of chat; they are a completely different product.
AI Billing Units Are Changing — The AWSization of AI Pricing
The billing unit shifts from “how many questions were asked” to “how much work the agent did.” This follows the exact same pattern as the transition from seat-based pricing in SaaS to usage-based pricing in the cloud. The AWS-ization of AI pricing has begun. This transition creates a much larger and more predictable revenue structure for AI infrastructure companies.
Token demand explodes even when the price drops.
The unit price of tokens has fallen to one-tenth of its level over the past two years, yet consumption has increased more than 100-fold. This mirrors the logic of the energy market, where consumption increases as prices drop. Goldman Sachs’ forecast of a 24x increase by 2030 is a long-term version of this logic. As prices drop, people consume more, and as consumption increases, the demand for infrastructure rises.
6Frequently Asked Questions
| question | answer |
|---|---|
| Will this Claude billing separation affect me as well? | If you have been using agent tools like Claude Code or OpenClaw via subscription, you will need to manage your usage after June 15. General users who use Claude.ai solely for chatting in their browsers are unlikely to notice any changes. |
| Is an agent really 100 times more expensive than chatting? | Based on Goldman Sachs estimates, always-on agents consume approximately 100,000 tokens per day, and chat consumes about 1,000 tokens per conversation. However, this is based on agents that are “always on,” and actual differences vary depending on the nature and configuration of the work. It is appropriate to understand this in terms of general trends rather than precise figures. |
| Should I buy Nvidia stock now? | This article is not an investment recommendation. However, the logical connection between an explosion in token consumption and increased demand for AI infrastructure is structural. It is noteworthy that the structure suggests Blackwell Ultra's cost-saving effects could further boost demand through the Jevons Paradox. Please examine valuation and macro risks separately. |
| Is it true that Anthropic has surpassed OpenAI's revenue? | Anthropic officially announced that it will surpass $30 billion in annualized revenue (ARR) on April 7, 2026. OpenAI is at the level of approximately $2.4 to $2.5 billion. However, OpenAI is refuting this, claiming that Anthropic's figures were overstated by about $8 billion due to differences in accounting methods, so an accurate comparison should be reserved until the IPO announcement. |
| If token prices continue to fall, won't the demand for AI infrastructure decrease as well? | On the contrary. While the unit price of tokens dropped to one-tenth over two years, consumption increased 100-fold. This is an AI version of the Jevons Paradox, which has been repeatedly confirmed in the energy market. Based on this dynamic, Goldman Sachs also projected a 24-fold growth by 2030. |

Conclusion — If the billing structure changes, the infrastructure map changes too
When Jensen Huang said, “A $500,000 engineer needs to spend $250,000 worth of tokens,” many people likely thought it was an exaggeration. However, the moment Claude’s pricing model separates the agent from the chat, you realize that this logic is being implemented on a real business model.
At this juncture, as AI transitions from chat apps to operational engines , the explosion in token consumption is not a prediction but a reality. Goldman Sachs' 24x forecast, Anthropic's ARR surpassing $30 billion, and Claude's pricing revamp are all pointing in the same direction.
It is time to add another lens through which to view AI. We must look not only at which model is smarter, but also at how the infrastructure running those models is monetized . And the biggest beneficiaries of this reality will be the companies that own the pipelines through which tokens flow .
All figures in this article are for informational purposes only and do not constitute investment advice. The Anthropic ARR of $30 billion is based on Anthropic's official announcement, and there are differing opinions within the industry regarding accounting methods. The Goldman Sachs 24x token forecast is a scenario analysis based on a report dated May 8, 2026. All investment decisions and responsibilities rest with the individual, and consulting with a professional financial advisor before making any significant decisions is recommended.
📌 Was this analysis helpful?
In the next post, we plan to cover “Token Cost Bomb — 3 Criteria for ROI Calculation Before Agent Implementation.”
We continue to track AI investment strategies in the agent era. Subscribe to notifications so you don't miss out.
How this content was produced
Aleph's research AI agent assisted with collecting and analyzing public data, creating charts and visuals, and structuring the draft. Davar personally reviewed and edited the sources, figures, reasoning, and final conclusions.
This content is for informational purposes only and is not personalized investment advice or an individual stock recommendation. Read the full disclaimer
© Aleph. All rights reserved.





