● Token-Cost Arms Race
The AI semiconductor race is no longer decided by “performance” but by “cost per token”
The core point in the AI semiconductor market is changing completely.
In the past, the question was “who has the faster GPU,” but now the real competitive edge is “how much can you lower the cost of generating one AI answer?”
In particular, as generative AI expands beyond chatbots into AI agents, coding automation, and enterprise workflow automation, token usage is exploding.
This trend is changing even the data center investment strategies of big tech companies such as Nvidia, AMD, Google, Amazon, Microsoft, Meta, and OpenAI.
In this article, we will cover in news style why the AI semiconductor race is shifting to “token economics,” and what Google’s Frozen v2, Cerebras Wafer Scale Engine, AMD Helios, Amazon Trainium, and Nvidia Vera Rubin strategies each mean.
1. The new standard in the AI industry: what is “cost per token”?
One of the most common phrases in the AI industry these days is “cost per 1 million tokens.”
Simply put, a token is the smallest unit of processing used when AI reads, understands, and generates a response.
Humans understand information around sentences and meaning, but AI breaks text into smaller pieces and processes them.
For example, even when you enter one question and receive one answer, hundreds to thousands of tokens may be used.
The issue is that each and every one of those tokens consumes electricity, memory, compute resources, and data center operating expenses.
- AI training cost: This is largely a one-time investment incurred when building the model.
- AI inference cost: This is a recurring cost that is incurred every time a user uses AI.
- Cost per token: This is a key indicator of whether an AI service can achieve long-term profitability.
In the end, from an AI company’s perspective, model performance alone is not enough.
It must generate better answers more cheaply, with less power, and faster.
This is why AI infrastructure and the semiconductor industry are being watched most closely in current global economic outlooks.
2. Why is token usage suddenly exploding?
The biggest reason is that the way AI is used is shifting from chatbots to AI agents.
In the past, chatbots had a structure where a user asked one question and AI provided one answer.
In that case, token usage was relatively limited.
But AI agents are different.
After receiving a user’s instruction, AI agents search on their own, write code, read files, call external tools, review the results, and then repeat the task again.
- Traditional chatbots: A simple structure of asking once and answering once.
- AI agents: They repeat multiple stages of work and consume new tokens each time.
- Coding agents: One task can consume hundreds of thousands to millions of tokens.
In particular, coding automation services and enterprise workflow automation services often run for long periods in the background.
That makes life easier for users, but from the AI company’s perspective, costs accumulate quickly.
This is also why some AI services impose weekly usage limits or consider usage-based billing once a certain threshold is exceeded.
3. The paradigm shift in AI semiconductor competition: from benchmarks to token economics
In the early days, AI semiconductor competition was relatively simple.
Companies that secured more GPUs, trained larger models, and posted higher benchmark scores were ahead.
But with the rise of the AI inference era, the math changed.
What matters now is how many tokens can be processed with the same power and the same cost.
This change is also having a major impact on data center investment strategies.
Rather than simply buying more GPUs, companies have begun looking at power efficiency, memory architecture, network bottlenecks, rack-level optimization, and even the software ecosystem as a whole.
The AI semiconductor market is now moving from “chip performance competition” to “system-wide cost competition.”
4. Two strategies for lowering cost per token
The ways the AI industry lowers cost per token can be broadly divided into two paths.
- The first is model optimization.
- The second is hardware optimization.
Model optimization involves activating only the necessary parts, compressing large models into smaller ones, or reusing repeated prompts.
Representative examples include MoE structures, model distillation, caching, and lightweight models.
Chinese DeepSeek and some AI companies promoting low-cost model strategies are drawing attention in this direction.
But the core point of this shift is hardware.
That is because if AI usage grows far beyond current levels, model optimization alone will hit its limits.
So big tech firms and semiconductor companies are redesigning AI semiconductors themselves around the cost of producing tokens.
5. Google Frozen v2: abandoning generality and optimizing for Gemini
Google’s strategy is among the most extreme.
Google is pushing the “Frozen v2” strategy, known as a dedicated chip optimized for its AI model Gemini.
The key is to give up general-purpose flexibility and dramatically improve token processing efficiency for a specific model.
Existing Google TPUs are closer to general-purpose AI accelerators like Nvidia GPUs, capable of running multiple AI models.
But Frozen v2 is different because it is a dedicated chip tailored to Gemini’s structure and data-processing method.
Based on the original text, Google expects token throughput per unit of power to improve by roughly 6x to 10x compared with the latest TPU.
- TPU: A general-purpose AI accelerator that can run multiple models.
- Frozen v2: A specialized AI semiconductor tailored for Gemini.
- Strategic direction: It prioritizes token efficiency over flexibility.
This strategy becomes even more important as Gemini is integrated across Google Search, YouTube, Cloud, Workspace, and other services.
As users increase, token costs explode, so Google is trying to design a semiconductor architecture suited to its own services from the start.
6. Cerebras Wafer Scale Engine: instead of connecting chips, make one huge one
Cerebras takes a completely different approach from traditional semiconductor design.
Usually, semiconductor companies connect multiple small chips to handle large-scale computation.
By contrast, Cerebras promotes a Wafer Scale Engine that uses an entire 12-inch wafer essentially as one giant chip.
The advantage of this approach is that it reduces the bottlenecks that occur when data moves between chips.
In AI inference, latency is extremely important.
Especially in services where users wait for real-time answers, the number of tokens processed per second determines perceived performance.
- Traditional approach: Multiple chips are connected for parallel processing.
- Cerebras approach: An entire wafer is used like one giant chip.
- Key advantage: It reduces chip-to-chip communication bottlenecks and latency.
The original text explains that Cerebras presented much higher tokens-per-second processing speed than Nvidia H-series clusters in frontier model inference.
However, since such figures can vary depending on model conditions, batch size, network environment, and software optimization, it is more important to look at the structural direction than at simple comparisons.
7. AMD Helios: competing at the rack level, not with a single GPU
AMD is emerging as the most realistic alternative for big tech companies trying to reduce dependence on Nvidia.
In particular, AMD Helios matters because it is not simply a GPU sale strategy, but a rack system integrating GPUs, CPUs, and networking into one unit.
Competition is now shifting from the performance of a single chip to total cost of ownership at the rack and data center level.
The reason major customers such as Microsoft, Meta, OpenAI, and Oracle are paying attention to AMD Helios is here.
Even if the initial purchase price is high, it can still be worth investing if long-term operating cost per token can be lowered.
- AMD Helios: A rack-level AI system integrating GPU, CPU, and networking.
- Key goal: Lower total cost of ownership and cost per token.
- Market meaning: It is an alternative that could crack Nvidia’s dominance.
Based on the original text, a single Helios unit may be more expensive than Nvidia’s competing product, but it is said to be attractive in terms of long-term operating cost.
AI infrastructure investment is now a long-term investment that includes not just hardware purchase, but also electricity, cooling, software transition costs, and operational automation.
8. Amazon Trainium and the self-developed chip race: why are big tech companies making their own chips?
Amazon is one of the most notable companies in the self-developed AI chip strategy.
AWS is trying to provide lower-cost AI computing options to cloud customers through Trainium and Inferentia series chips.
The original text explains that Amazon Trainium 3 aims for a structure that is cheaper than Nvidia GPU-based cloud solutions in terms of cost to process 1 million tokens.
The reason big tech companies make their own chips is clear.
If they keep buying Nvidia GPUs, performance is excellent, but they have to accept margin pressure and supply constraints.
On the other hand, making their own chips allows optimization for their own services and can improve cloud profitability over the long term.
- Amazon: Strengthens cloud cost competitiveness with its own AI chips for AWS.
- Microsoft: Tries to improve the efficiency of its own AI infrastructure through the Maia series of chips.
- Meta: Optimizes recommendation, advertising, and generative AI workloads through its own AI accelerators and cooperation with Broadcom.
- OpenAI: Is pursuing a long-term self-developed chip strategy to lower inference costs compared with Nvidia GPUs.
This trend shows a structural change in the AI semiconductor market.
In the past, semiconductor companies made chips and cloud companies bought them, but now cloud companies are directly entering semiconductor design.
9. Is the Nvidia era over? Not yet
Although the move away from Nvidia is growing stronger, it is hard to say that the Nvidia era is ending immediately.
Nvidia still holds overwhelming market share in the AI accelerator market.
Above all, Nvidia’s real strength is not just the GPU itself, but the CUDA software ecosystem.
AI developers and companies have already built workflows based on CUDA to train, optimize, and deploy models.
This software lock-in effect is stronger than many expect.
Even if the chip is expensive, the ability to run existing code easily, the rich developer ecosystem, and the abundance of troubleshooting resources are major advantages.
- Nvidia’s strengths: GPU performance, CUDA ecosystem, developer base, supply chain, and system integration capabilities.
- Nvidia’s weaknesses: High prices, supply constraints, and dependency risk for big tech companies.
- Competitors’ opportunity: Offering lower cost per token for specific workloads.
Nvidia knows this trend too.
That is why it is trying to improve inference efficiency at the entire system level rather than just at the chip level through structures like the next-generation rack system Vera Rubin NVL72.
In the end, Nvidia is also moving toward “cheaper tokens,” not just “faster GPUs.”
10. Why do people say token costs fall even when buying expensive chips?
This may be confusing for many people.
The question “If the equipment is more expensive, how can costs go down?” naturally comes up.
The key is to distinguish between initial purchase price and long-term operating cost.
In AI data centers, server price is not the only thing that matters.
Electricity, cooling, rack space, network costs, failure response, software optimization, and labor costs all need to be included.
Even if the initial equipment price is high, it can be cheaper in the long run if it processes more tokens with the same amount of power.
- Initial cost: The amount paid when purchasing chips and servers.
- Operating cost: Electricity, cooling, maintenance, and software operating expenses.
- Cost per token: The true competitiveness that reflects both initial and operating costs.
From this perspective, it becomes understandable why big tech companies review alternative chips or systems that appear more expensive than Nvidia’s.
They are looking not at today’s equipment price, but at the cost of running AI services for the next 3 to 5 years.
11. The real core point that other news often misses
The most important part of this AI semiconductor race is not “whether Nvidia wins or loses.”
The real core point is that the AI industry’s profit model is being reorganized around cost per token.
- First, cost structure is becoming more important than revenue for AI companies.
- Second, the spread of AI agents is highly likely to increase token usage exponentially.
- Third, big tech’s self-developed chip efforts are not just technology bragging but a margin defense strategy.
- Fourth, the ability to secure data center power is emerging as a key variable in AI competitiveness.
- Fifth, future AI service pricing is likely to move from subscriptions to usage-based billing.
From an investment perspective, looking only at AI semiconductor companies is not enough.
You also need to look at power infrastructure, cooling technology, networking equipment, memory semiconductors, and cloud operating efficiency.
As the AI industry grows, it is not just semiconductors that grow; the entire data center investment value chain moves as well.
12. What to watch next: who can generate intelligence more cheaply?
In the future AI market, the winner may not simply be the company with the smartest model.
The company that can run the smartest model the cheapest is more likely to win.
Goldman Sachs projects that global AI token usage will increase significantly over the long term compared with current levels.
If AI agents begin to be widely used in enterprise operations, software development, customer support, financial analysis, healthcare, and manufacturing automation, token demand could grow beyond comparison with today.
In this situation, companies need to answer three questions.
- How much does our AI service cost per token?
- Will profitability hold if usage grows 10 times?
- Can we secure stable AI infrastructure while reducing dependence on Nvidia?
In the end, the AI semiconductor war is moving from a performance race to a cost competition.
And this change is likely to connect the profitability of the generative AI industry, cloud pricing policies, global data center investment, semiconductor supply chains, and big tech stock trends.
< Summary >
The core point of the AI semiconductor race is now cost per token, not performance.
As AI agents spread, token usage is exploding, and inference cost is determining the profitability of AI companies.
Google is trying to reduce generality and increase efficiency with Gemini-optimized chips.
Cerebras has chosen a strategy of reducing bottlenecks with wafer-scale chips.
AMD has entered rack-level total cost of ownership competition through Helios.
Amazon, Microsoft, Meta, and OpenAI are trying to reduce dependence on Nvidia with their own AI chips.
But Nvidia still maintains a strong position with the CUDA ecosystem and next-generation rack systems.
The future winner in the AI industry is likely to be the company that produces intelligence more cheaply, not the company that has the faster chip.
[Related Articles…]
- How the AI Semiconductor Market Is Changing and Big Tech’s Investment Strategy
- An Analysis of the AI Infrastructure Competition After Nvidia
*Source: [ 티타임즈TV ]
– AI 칩 경쟁의 패러다임이 ‘성능’에서 ‘원가’로 바뀌고 있다


