Goldman Sachs Asia Pacific Internet team analyst Ronald Keung on Monday (3rd) raised his forecast for the combined annual recurring revenue (ARR) of Chinese large model vendors in 2024 from $10 billion to $13 billion, and projected it to climb to $125 billion by 2030—equivalent to a roughly 25-fold expansion over five years.

Jefferies noted the same day, based on OpenRouter data, that Chinese models have ranked among the top five globally in API call volume for 14 consecutive weeks, with DeepSeek V4 Flash consuming 7.22 trillion tokens in a single week to claim the top spot.

According to Goldman Sachs, China's large model API and subscription revenue is expected to grow from an estimated RMB 35 billion in 2026 to RMB 879 billion in 2030. Daily token consumption is projected to surge from 350 trillion to 4,600 trillion, reflecting a 25-fold annual growth over five years with no signs of slowing. Structurally, a 'two-tier market' is emerging. The high-end segment, led by Zhipu's GLM-5.2 and Alibaba's Qwen3.7 Max, prices one million tokens at around $1, capturing an intelligence premium. The low-end segment, targeting Agent tasks, has prices slashed to $0.06–$0.20, leveraging economies of scale.

Goldman Sachs revised its revenue estimates for Zhipu and MiniMax upward by 35% and 22–64% respectively, signaling a repricing of major vendors' commercialization capabilities.

However, Goldman notes that total training and inference costs in 2026 will reach $11 billion—exceeding ARR at that time—with EBIT still negative. API gross margins remain at 20–30%, with the profitability inflection point pegged at 2030, when EBIT margins are expected to reach 18% and total profit $23 billion—indicating a 'burn first, monetize later' model, but with a clearly defined endpoint.

The real reason Goldman changed its stance is the shift in token consumption from 'humans' to 'Agents.' For the same programming task, automated Agent workflows consume an average of 4.17 million tokens, compared to just 3,390 for regular Q&A—over a thousandfold difference.

According to data disclosed by OpenCode, DeepSeek V4 Flash processes 8 trillion tokens per day, and Agent applications like Hermes have surpassed human chat to become the largest consumption segment on OpenRouter.

Goldman points out that Agent and coding scenarios simultaneously demand models that are 'smart enough and cheap enough.' Chinese models, leveraging MoE sparse activation and cache hit discounts (DeepSeek charges only 2% of list price for hits), have reduced the daily 'fuel cost' of top U.S. models from $52–$104 to just $13—a 4–8x price advantage that scales into hundreds of times cost difference in large deployments.

Goldman also defines 2025 as the 'DeepSeek Moment' (cost efficiency), and 2024 as the 'Zhipu Moment' (GLM-5.2 joining the global top tier in code and Agent capabilities). The performance gap between top U.S. and Chinese models has narrowed to about 2.7%. MiniMax's H3 ranks first in video editing AA benchmarks, and ByteDance's Seedance leads in multimodal AI, proving Chinese models are favored not just for being cheap.

Cloud spillover effects are also beginning. Goldman expects Alibaba Cloud's AI revenue share to approach 50%, with AI products achieving nine consecutive quarters of triple-digit growth. Cloud service providers (CSPs) are shifting from selling compute power to a hybrid MaaS + IaaS model.

Goldman concludes that the competitive landscape of China's large models has been rewritten. The core is no longer parameters or benchmarks, but who can minimize inference costs the most, capture the most Agent/coding tokens, and fastest convert to cloud revenue. The shift from follower to rule-maker is no longer a prediction—it's a reality already evident in Q2 2024 data.

FACT BOX

  • Source: PR Times
  • Category: Survey
  • Organizations: MiniMax / DeepSeek / OpenCode
  • Products / services: AI API