The multi-model aggregation platform OpenRouter has released its global AI large model usage weekly ranking for the period from July 27 to August 2, showing that DeepSeek V4 Flash achieved first place with 7.22 trillion tokens in usage.

Additionally, according to a post by the open-source project team OpenCode, the usage volume of DeepSeek V4 Flash on their platform surged dramatically, reaching 8 trillion tokens processed in a single day on August 1. Of this, 5 trillion tokens were consumed through free trial quotas, while 3 trillion tokens were used by developers paying via the OpenCode platform.

According to a report by Jiemian News on Wednesday the 5th, OpenRouter data shows that last week's total global AI large model usage reached 56.8 trillion tokens, a 2.07% decrease from the previous week.

Among the listed AI large models, Chinese AI models recorded a weekly usage of 28.13 trillion tokens, down 14.76%. In contrast, U.S. AI models saw a weekly usage of 4.38 trillion tokens, a sharp increase of 87.18%. Chinese models have now led U.S. models in weekly usage for fourteen consecutive weeks, firmly holding the top global position.

In OpenRouter's current weekly ranking, DeepSeek V4 Flash 0423 remains in first place with a usage volume of 6.92 trillion tokens.

This model was first launched in April this year, emphasizing a balance between inference speed, long-context capability, and cost efficiency, primarily offering high-throughput inference services to developers and commercial enterprises.

In second place is Xiaomi's (01810-HK) MiMo-V2.5, with a weekly usage of 5.1 trillion tokens, a sharp 52% decline from the previous week, indicating a cooling in market enthusiasm.

Xiaomi's model officially entered public beta on April 23 and completed full open-sourcing of its series by the end of April. It adopts a Mixture-of-Experts (MoE) architecture, with total parameters exceeding the trillion level and 42 billion active parameters. It comes standard with a 1 million-token ultra-long context window and supports multimodal interactions across text, speech, and images.

In third place is Tencent's (00700-HK) self-developed large model HunYuan Hy3, with a usage scale of 5.01 trillion tokens. Hy3 was the fastest-growing model during the week including July 26, with a weekly usage of 3.94 trillion tokens at that time, an increase of over 999%.

HunYuan Hy3 was officially open-sourced on July 6, with a total of 295 billion parameters and only 21 billion activated per inference. It adopts a fast-slow thinking fusion architecture, supporting up to 256K context, with significant improvements in code generation and intelligent interaction capabilities.

FACT BOX

  • Source: PR Times
  • Category: New Product
  • Organizations: OpenRouter / OpenCode
  • Products / services: DeepSeek V4 Flash / MiMo-V2.5