Skip to content
AI news, tool reviews, expert columns, prompts, agents and practical automation workflows.
News

Chinese AI labs shift focus from small models to trillion-parameter giants as Qwen 4 training begins

Alibaba, DeepSeek and Moonshot AI are building models with trillions of parameters, while the company behind Qwen confirmed on September 22 that Qwen 4 is in training and its roadmap targets 5 to 10 trillion parameters for future versions.

News Published 26 September 2026 4 min read Maya Turner
Chinese AI labs are now building models with trillions of total parameters, including Alibaba's Qwen3.8-Max at 2.4 trillion and Moonshot AI's Kimi K3 at 2.8 trillion, with Qwen 4
Imagen destacada del articulo fuente

Chinese AI companies that gained attention for compact, device-ready models are now racing in the opposite direction: building some of the largest language models ever disclosed. The shift became clearer this week as Alibaba confirmed that Qwen 4 is already in training and that its roadmap for Qwen 4.5 and Qwen 5 targets models between 5 and 10 trillion total parameters.

The move marks a strategic pivot for a country that had been associated primarily with efficient, smaller-scale architectures such as Qwen’s compact variants and MiniCPM. Now, several major Chinese labs are publishing parameter counts that rival or exceed anything publicly acknowledged by US counterparts.

Key facts

Model Total parameters Active parameters per step Status
DeepSeek-V3 671 billion Not disclosed Released
Kimi K2 (Moonshot AI) 1 trillion Not disclosed Released
Qwen3.8-Max (Alibaba) 4 trillion ~95 billion Released
Kimi K3 (Moonshot AI) 8 trillion Not disclosed Released
Qwen 4 (Alibaba) Not disclosed Not disclosed In training (confirmed Sep 22)
Qwen 4.5 / Qwen 5 (roadmap) 5–10 trillion Not disclosed Planned

What the parameter numbers actually mean

Total parameter counts have become a headline metric, but they do not directly translate to capability. Most of these large-scale models use a Mixture of Experts (MoE) architecture, which activates only a subset of parameters for each input. SemiAnalysis pointed to this design in GPT-4 as early as 2023, though OpenAI never confirmed it publicly.

Alibaba’s Qwen3.8-Max illustrates the gap: it has 2.4 trillion total parameters but activates roughly 95 billion per processing step. That means the model can store broad knowledge across many specialized sub-networks while keeping inference costs closer to those of a much smaller model.

The distinction matters for developers and enterprises evaluating these models. A trillion-parameter total does not necessarily mean trillion-parameter inference costs, but it does imply larger memory footprints for serving and more complex infrastructure requirements.

Scaling alone is not a guarantee of quality

Parameter count is one variable among several. DeepMind demonstrated in 2022 with its Chinchilla paper that a 70-billion-parameter model could outperform a 280-billion-parameter model (Gopher) across most benchmarks when trained on roughly four times more data under the same compute budget. Architecture, data quality, training compute and dataset size all influence final performance.

No public benchmarks yet compare the latest Chinese trillion-parameter models against each other or against US counterparts such as OpenAI’s GPT-5.6 Sol or Anthropic’s Claude Fable 5.1, neither of which have disclosed parameter counts. Direct comparison is not possible from published specifications alone.

The 10-trillion-parameter frontier

Alibaba confirmed on September 22 that Qwen 4 is in training and that its published roadmap includes Qwen 4.5 and Qwen 5 at scales of 5 to 10 trillion total parameters. ByteDance is reportedly pursuing a similar scale, according to Financial Times, though the company has not confirmed this publicly.

These numbers would represent a step change even from the current generation. Kimi K3 at 2.8 trillion and Qwen3.8-Max at 2.4 trillion already exceed the trillion-parameter mark that was rare in 2025. A 10-trillion-parameter model would require substantially more training compute and serving infrastructure, and it remains unclear whether the performance gains will justify the cost.

Small models and large models serve different use cases

The Chinese AI ecosystem continues to develop both ends of the size spectrum. Compact models such as Qwen’s smaller variants and MiniCPM are designed for on-device inference, reduced memory requirements and lower operational costs. Large models target broader capability across more demanding and varied tasks, at the cost of greater infrastructure needs.

This dual-track approach means Chinese labs are not abandoning the small-model strategy that brought them attention. Instead, they are investing in both extremes: efficient models for edge deployment and colossal systems for cloud-based reasoning and research.

Source: Xataka IA, “Las empresas chinas lleva la delantera en modelos ‘pequeños’ de IA. Ahora lo que quieren es justo lo contrario: modelos colosales,” by Javier Marquez, published September 25, 2026. https://www.xataka.com/robotica-e-ia/empresas-chinas-lleva-delantera-modelos-pequenos-ia-quieren-justo-contrario-modelos-colosales

Source

Xataka IA Publicacion original: 2026-09-25T21:48:17+00:00