Chinese AI labs shift focus from small models to trillion-parameter giants as Qwen 4 training begins
Alibaba, DeepSeek and Moonshot AI are building models with trillions of parameters, while the company behind Qwen confirmed on September 22 that Qwen 4 is in training and its roadmap targets 5 to 10 trillion parameters for future versions.


Chinese AI companies that gained attention for compact, device-ready models are now racing in the opposite direction: building some of the largest language models ever disclosed. The shift became clearer this week as Alibaba confirmed that Qwen 4 is already in training and that its roadmap for Qwen 4.5 and Qwen 5 targets models between 5 and 10 trillion total parameters.
The move marks a strategic pivot for a country that had been associated primarily with efficient, smaller-scale architectures such as Qwen’s compact variants and MiniCPM. Now, several major Chinese labs are publishing parameter counts that rival or exceed anything publicly acknowledged by US counterparts.
Key facts
| Model | Total parameters | Active parameters per step | Status |
|---|---|---|---|
| DeepSeek-V3 | 671 billion | Not disclosed | Released |
| Kimi K2 (Moonshot AI) | 1 trillion | Not disclosed | Released |
| Qwen3.8-Max (Alibaba) | 4 trillion | ~95 billion | Released |
| Kimi K3 (Moonshot AI) | 8 trillion | Not disclosed | Released |
| Qwen 4 (Alibaba) | Not disclosed | Not disclosed | In training (confirmed Sep 22) |
| Qwen 4.5 / Qwen 5 (roadmap) | 5–10 trillion | Not disclosed | Planned |
What the parameter numbers actually mean
Total parameter counts have become a headline metric, but they do not directly translate to capability. Most of these large-scale models use a Mixture of Experts (MoE) architecture, which activates only a subset of parameters for each input. SemiAnalysis pointed to this design in GPT-4 as early as 2023, though OpenAI never confirmed it publicly.
Alibaba’s Qwen3.8-Max illustrates the gap: it has 2.4 trillion total parameters but activates roughly 95 billion per processing step. That means the model can store broad knowledge across many specialized sub-networks while keeping inference costs closer to those of a much smaller model.
The distinction matters for developers and enterprises evaluating these models. A trillion-parameter total does not necessarily mean trillion-parameter inference costs, but it does imply larger memory footprints for serving and more complex infrastructure requirements.
Scaling alone is not a guarantee of quality
Parameter count is one variable among several. DeepMind demonstrated in 2022 with its Chinchilla paper that a 70-billion-parameter model could outperform a 280-billion-parameter model (Gopher) across most benchmarks when trained on roughly four times more data under the same compute budget. Architecture, data quality, training compute and dataset size all influence final performance.
No public benchmarks yet compare the latest Chinese trillion-parameter models against each other or against US counterparts such as OpenAI’s GPT-5.6 Sol or Anthropic’s Claude Fable 5.1, neither of which have disclosed parameter counts. Direct comparison is not possible from published specifications alone.
The 10-trillion-parameter frontier
Alibaba confirmed on September 22 that Qwen 4 is in training and that its published roadmap includes Qwen 4.5 and Qwen 5 at scales of 5 to 10 trillion total parameters. ByteDance is reportedly pursuing a similar scale, according to Financial Times, though the company has not confirmed this publicly.
These numbers would represent a step change even from the current generation. Kimi K3 at 2.8 trillion and Qwen3.8-Max at 2.4 trillion already exceed the trillion-parameter mark that was rare in 2025. A 10-trillion-parameter model would require substantially more training compute and serving infrastructure, and it remains unclear whether the performance gains will justify the cost.
Small models and large models serve different use cases
The Chinese AI ecosystem continues to develop both ends of the size spectrum. Compact models such as Qwen’s smaller variants and MiniCPM are designed for on-device inference, reduced memory requirements and lower operational costs. Large models target broader capability across more demanding and varied tasks, at the cost of greater infrastructure needs.
This dual-track approach means Chinese labs are not abandoning the small-model strategy that brought them attention. Instead, they are investing in both extremes: efficient models for edge deployment and colossal systems for cloud-based reasoning and research.
Source: Xataka IA, “Las empresas chinas lleva la delantera en modelos ‘pequeños’ de IA. Ahora lo que quieren es justo lo contrario: modelos colosales,” by Javier Marquez, published September 25, 2026. https://www.xataka.com/robotica-e-ia/empresas-chinas-lleva-delantera-modelos-pequenos-ia-quieren-justo-contrario-modelos-colosales
Source
Xataka IA Publicacion original: 2026-09-25T21:48:17+00:00
Maya Turner
Colaborador editorial.
