The next war in AI is not about bigger models, but higher-efficiency intelligence—redefining the global AI industry’s next decade through Token Economics and Edge AI.

For the past three years, the global AI industry has revolved around one question: Who owns the largest model? Who built the most GPU clusters? Who invested the most capital in constructing massive data centers?

From parameter scale and training compute to infrastructure, AI seemed to be an endless race of "bigger is better." However, if we closely observe changes in the global AI market over the past six months, we’ll find that what’s truly shifting the industry’s direction is no longer just model capability—but a more practical question: Who can perform the most AI inference at the lowest cost?

This question may appear to be merely about cost control, but it could determine the landscape of the AI industry for the next decade.

AI Returns to Business Fundamentals

When generative AI first emerged, people were amazed by the creativity and reasoning capabilities of large models. The market chased model leaderboards—who broke the latest benchmark, who launched a model with even more parameters. But once inside enterprises, managers saw a different report—not model scores, but ever-rising monthly AI bills.

When tens of thousands of employees use AI daily to write emails, organize documents, search knowledge, create presentations, respond to customer service queries, or simply complete routine administrative tasks, every token consumed directly impacts operational costs.

For the first time, AI has shifted from a technical issue to a financial one. Enterprises are realizing that what’s truly expensive isn’t adopting AI—it’s the ongoing inference cost (Inference Cost) incurred every day.

Thus, a new industry trend is forming. Companies no longer ask: "Which model is the most powerful?" Instead, they’re asking: "Which model best fits my cost structure?"

Token Economics Is Reshaping the AI Market

Over recent months, usage data released by various AI platforms have revealed the same signal. The fastest-growing models aren’t necessarily the most capable or highest-priced closed models, but rather those that are significantly cheaper, easier to deploy, and more open-source.

This indicates that the AI industry is gradually shifting from "Model Competition" to "Token Economics."

Enterprises are now calculating the cost of each token, just as they previously managed cloud resources, network bandwidth, or electricity costs. Because what businesses truly seek isn’t the most expensive intelligence, but intelligence with the highest return on investment (ROI of Intelligence).

We can broadly categorize AI work into two types:

First, high-value tasks such as strategic planning, complex reasoning, R&D innovation, and multi-step analysis—these still require the most advanced large models.

Second, repetitive daily operations such as document organization, customer service responses, equipment inspections, image recognition, process automation, administrative tasks, and knowledge searches.

These tasks make up the majority of enterprise AI usage, yet don’t necessarily require the most expensive models. No company would rent an entire supercomputer just to generate an Excel report. Likewise, no enterprise wants to use the highest-cost AI inference for every email or customer interaction.

Therefore, more and more companies are building tiered AI (Tiered AI) architectures—assigning complex tasks to large models and standardized work to more efficient, lower-cost models. AI is being finely configured like enterprise ERP systems, not treated as a one-size-fits-all solution.

AI Enters the True Era of Economics

The focus of AI competition is shifting. The real competition in AI is moving from model development to inference efficiency.

The first stage was about who could train the largest model.

The second stage is about who can make each inference cheaper.

Going forward, the truly important metrics won’t just be model capability, but cost per token (Cost per Token), such as:

- Inferences per watt of power (Inference per Watt) - Business value created per dollar of computing power (Business Value per Dollar)

In other words, AI is entering a true era of economics.

Cloud AI Won’t Disappear, But Edge AI Will Rise Rapidly

Many mistakenly assume that the importance of large data centers will decline. The opposite is true. Large data centers will continue to play irreplaceable roles—they train world-class models, provide high-end inference capabilities, manage global knowledge, and support the entire AI ecosystem.

What’s truly changing is how AI is deployed.

Future AI will adopt a hybrid "Cloud + Edge" architecture. The cloud creates intelligence (Create Intelligence), while edge devices consume intelligence (Consume Intelligence).

Just as no one today sends every smartphone computation back to a data center, billions of future AI devices cannot rely on the cloud for every identification or decision. Real-time performance, energy efficiency, network bandwidth, data privacy, and system resilience all demand that AI be closer to where data is generated.

Thus, we’ll see AI evolve from Cloud Native to Cloud-to-Edge Native.

NPU Will Redefine AI Inference

This is why I believe the most important AI chip over the next decade won’t just be the GPU. GPUs revolutionized AI model training. NPUs (Neural Processing Units) will transform AI inference. They aren’t competitors, but different roles within the AI ecosystem. GPUs handle high-performance computing and drive model breakthroughs. NPUs specialize in low-power, low-latency, high-efficiency inference, enabling AI deployment in cars, robots, factories, autonomous drones, smart cameras, medical devices, and various end-user devices.

The future world won’t consist of only hundreds of massive data centers. What truly transforms the world will be billions of intelligent devices with AI capabilities. If AI exists only in the cloud, it will never truly enter people’s lives. Only when AI can compute, decide, and learn instantly on devices does it become true infrastructure—not just a cloud service.

The Next AI Revolution Is About Efficiency

In AI’s first decade, the race was about who could build the largest model. In its second decade, it’s about who can deploy intelligence globally at the lowest cost and highest efficiency.

What truly determines industry leadership is no longer just model parameters, but overall computing architecture; not just the number of GPUs, but the collaborative efficiency of the entire Cloud-to-Edge ecosystem; not just data center scale, but the value created per watt of energy, per token, and per inference.

From large models to open-source models, from centralized computing to edge intelligence, from pursuing peak performance to optimizing cost-effectiveness, the AI industry is transitioning from a technological race to a revolution centered on efficiency, commercialization, and widespread adoption.

The winners of the next decade won’t necessarily be those with the largest models, but those who enable billions of devices worldwide to use AI securely, reliably, and at the lowest possible cost.

And when AI truly becomes as ubiquitous as electricity and the internet, Edge AI and NPUs will no longer be mere semiconductor buzzwords—they will become the core engines driving the next wave of global digital economic growth.

*Author: Liu Chun-Cheng, Founder & CEO of Kneron, Senior Policy Advisor at Xinyou Association This article is provided and authorized for publication by Xinyou Association

FACT BOX

  • Source: PR Times
  • Category: News
  • Organizations: Apple / Google
  • Products / services: NPU / Edge AI