Alibaba Zhenwu V900: AI chip faces the scaling test

By Julien Mercier

an hour ago


Centre de données asiatique au crépuscule, avec serveurs, refroidissement et ingénieurs, illustrant l’intégration entre puces IA, cloud et énergie.
Asian data center at dusk, with servers, cooling and human supervision, illustrating the integration of AI chips, cloud infrastructure and energy. Credits: Nezna/generated by IA.
In short
  • Alibaba presents the T-Head-designed Zhenwu V900 as roughly three times faster than the previous M890 and says clusters can scale to as many as 500,000 accelerators.
  • Mass production is planned for the first quarter of 2027: no independent, reproducible comparison with Nvidia, Huawei or other accelerators is publicly available yet.
  • Alibaba is pairing the chip roadmap with future Qwen models containing 5 trillion to 10 trillion parameters and a target of more than 20 GW of global data-center capacity by 2032.
  • The central challenge is not only chip performance but whether Alibaba can manufacture, connect, power, cool and operate these accelerators at scale and at a useful cost for customers.

Alibaba used its September 22, 2026 Apsara conference in Hangzhou to present an ambition that goes beyond launching another processor. The group wants to advance its Qwen models, T-Head-designed accelerators, interconnect technology and data centers at the same time. At the center of that strategy is the Zhenwu V900, its next-generation chip for artificial-intelligence training and inference.

Reuters, Associated Press, China's The Paper and Seoul Economic Daily converge on the main figures disclosed at the event: Alibaba says the V900 delivers about three times the performance of the Zhenwu M890, can support clusters of up to 500,000 accelerators, will enter mass production in the first quarter of 2027 and is part of a plan to exceed 20 GW of data-center capacity by 2032. That convergence confirms the announcements were made; it does not independently verify the technical claims.

Three times the performance, but three times what?

The first limitation is the headline threefold improvement. In the sources reviewed, Alibaba does not provide enough information to establish exactly what the multiplier measures: compute throughput at a given numerical precision, training, inference, tokens per second, energy efficiency or complete-system performance. Without a comparable workload, protocol and power envelope, the figure remains a vendor-supplied measurement.

The previous generation provides a useful contrast. When Alibaba introduced the Zhenwu M890 in May 2026, it disclosed more technical specifications: 144 GB of memory, 800 GB/s of inter-chip bandwidth, support for numerical formats from FP32 to FP4 and a proprietary ICN interconnect. Alibaba also announced an ICN switching chip providing up to 25.6 Tbps of aggregate bandwidth and full-bandwidth interconnection across 64 accelerators.

For the V900, the information released on September 22 focuses mainly on performance and scaling targets. The absence, so far, of detailed memory capacity, memory bandwidth, power draw, performance at different numerical precisions and manufacturing-process data makes a rigorous comparison with accelerators from Nvidia, Huawei, AMD, Cambricon or other suppliers impossible.

The same caution applies to the description of the V900 as China's most powerful AI chip, wording used around the launch and repeated in several reports. None of the independent sources reviewed provides a standardized ranking that would verify that claim.

500,000 chips: the bottleneck moves toward infrastructure

The most striking figure may be the proposed 500,000-accelerator cluster. It does not mean Alibaba operates such a system today. Reuters says mass production is scheduled for the first quarter of 2027, so 500,000 describes an announced architectural scaling ceiling, not an observed deployment.

At that scale, adding processors is not enough. Useful performance depends on the network connecting them, memory, storage, node synchronization, failure rates, distributed software, electricity supply and cooling. The engineering problem increasingly becomes one of preventing those layers from becoming more restrictive than the accelerator itself.

The M890 generation shows Alibaba is already working at this system level. Its Panjiu AL128 server places 128 accelerators in a rack, with Alibaba claiming internal bandwidth at the petabyte-per-second scale. Moving from dozens or hundreds of tightly connected accelerators to hundreds of thousands nevertheless requires several networking layers and much more complex orchestration. In the sources examined, Alibaba has not publicly detailed the topology that would connect 500,000 V900 cards or the efficiency that could be maintained close to that scale.

Large AI cluster with accelerator racks, optical interconnects, liquid cooling, electrical distribution and engineers supervising the infrastructure.
At very large scale, AI-system performance depends as much on interconnects, memory, power and cooling as on the accelerators themselves. Credits: Nezna/generated by IA.

A scale shift that can be compared with volumes already reached

In May, Alibaba said it had cumulatively delivered more than 560,000 Zhenwu chips and had more than 400 external customers across about 20 industries. Those figures come directly from Alibaba and do not constitute an independent audit, but they provide a useful order of magnitude: the proposed maximum of 500,000 V900s in one cluster is close to the number of Zhenwu units Alibaba said it had cumulatively delivered across previous generations and customers at that point.

That does not make the project unachievable. It illustrates the scale change being proposed. The V900 challenge is as industrial as it is microelectronic: enough chips have to be manufactured, memory and networking equipment must be available in volume, infrastructure has to be built, and acceptable efficiency must be maintained when thousands of nodes operate together.

Associated Press cites Counterpoint Research analysts who make a related point: progress by Chinese chip designers also depends on the capabilities of the foundries that physically manufacture their processors. The sources reviewed do not establish the V900's fabrication process or foundry with sufficient confidence. A precise claim on either point would therefore be premature.

Up to 10 trillion parameters: bigger does not automatically mean better

The V900 accompanies another major announcement. Alibaba plans future Qwen models with between 5 trillion and 10 trillion parameters. A comparison with its current flagship shows why total parameter counts have to be interpreted cautiously.

Alibaba Cloud's official documentation states that Qwen3.8-2.4T-A95B, launched in August 2026, uses a sparse Mixture-of-Experts architecture containing 2.4 trillion total parameters, of which approximately 95 billion are activated at each step. Total parameter count therefore describes the model's full parameter set, but not directly the amount of computation required for every token.

Alibaba has not yet publicly detailed the architecture of the planned 5-trillion to 10-trillion-parameter models. It therefore cannot be assumed that a model four times larger will require exactly four times more compute or provide four times more capability. Parameter count also does not directly measure reliability, answer quality or economic efficiency. Training data, architecture, active parameter count, post-training, available tools and the ability to complete an end-to-end task with few errors remain critical.

20 GW by 2032 makes energy part of the product

Alibaba is simultaneously targeting more than 20 GW of worldwide data-center capacity operated by Alibaba Cloud by 2032. Reuters and The Paper also report Eddie Wu acknowledging that shortages across the AI data-center supply chain are already limiting how quickly the company can expand computing capacity.

The 20 GW number itself requires caution. Published information does not provide a sufficiently precise technical definition to determine whether it maps to usable IT load, site electrical capacity or a broader metric. It would therefore be misleading to translate it directly into a number of V900 accelerators or a quantity of models that could be trained.

It nevertheless shows how the AI equation is moving into physical infrastructure. Electricity access, transformers, cooling, networking, data-center construction, memory and component availability are becoming as strategic as processor design.

A strategy extending far beyond model research

Vertical integration can give Alibaba more control over costs and over the fit between models, accelerators and software. An internal chip can reduce reliance on some third-party suppliers and allow the group to optimize hardware around Alibaba Cloud workloads. But it also transfers more risk to Alibaba: silicon design, manufacturing, interconnects, software tooling, maintenance, energy procurement and data-center depreciation.

Alibaba announced in 2025 that it would spend at least 380 billion yuan, roughly $53 billion, over three years on cloud and AI infrastructure. The scale of that commitment shows that competition is no longer only about which model ranks highest on a benchmark, but about the ability to supply computing capacity reliably and at large scale.

For customers, the V900's success will therefore be measured less by a theoretical maximum than by the cost of a successfully completed task, latency, service availability, error rates and energy consumption. Those metrics matter particularly for continuous workloads such as assistants, software agents, document analysis and multimodal systems.

One announcement, different regional readings

Reuters primarily frames the announcement as Alibaba building a complete technology stack spanning models, semiconductors and data centers. It also places the V900 in the context of US restrictions on advanced technologies while explicitly attributing the performance claims to Alibaba.

Associated Press places greater emphasis on US-China technology competition. Its most useful external contribution comes from the Counterpoint analysts it cites, who note that manufacturing capability can matter as much as chip design itself.

The Paper, based in mainland China, devotes more attention to Eddie Wu's industrial argument: rising token demand, deployment of M890 systems, supply bottlenecks and the construction of a global compute network. This framing helps explain Alibaba's strategy without reducing the subject to a contest with Nvidia.

Seoul Economic Daily, based in South Korea, takes a more competitive and financial angle and explicitly presents the V900 as a challenge to Nvidia. The publication states, however, that some information comes from Bloomberg and that its English version is AI-translated. That transparency is a reason to corroborate its strongest wording rather than rely on it in isolation.

Alibaba Cloud is used here as a primary source for M890 technical specifications, previous-generation shipments and Qwen3.8 architecture. The conflict of interest is direct: Alibaba is documenting its own products. Its disclosures establish what the company says it has designed or deployed, not independent validation of performance.

Reducing dependence on Nvidia does not replace its ecosystem

The V900 could reduce strategic reliance on foreign accelerators if Alibaba can manufacture enough units and provide a software environment that meets developers' needs. But owning a processor is not the same as replacing an ecosystem. Libraries, compilers, debugging tools, distributed frameworks and established engineering practices create dependencies of their own.

Alibaba's SAIL software stack and proprietary ICN interconnect show that it is treating the challenge as a system rather than simply a chip. That could enable close optimization between its models and infrastructure, but may also create compatibility costs for customers accustomed to other platforms.

It is therefore too early to describe the V900 as a general substitute for Nvidia GPUs. A more precise conclusion is that Alibaba is trying to build a proprietary compute path capable of handling a growing share of its own AI workloads. How far that substitution can go will depend on measured performance, available volumes and software maturity.

What remains genuinely unknown

Not independently verified: the roughly threefold V900 performance improvement over the M890, its description as China's most powerful AI chip and effective scaling to a 500,000-accelerator cluster originate with Alibaba. No independent public benchmark identified in the sources reviewed currently confirms them.

Not disclosed with sufficient precision: manufacturing process, foundry, power draw, memory capacity and bandwidth, performance at different numerical precisions, performance per watt, V900 pricing and real-world cluster efficiency at different scales.

Future targets rather than achieved results: first-quarter 2027 mass production, 5-trillion to 10-trillion-parameter models and more than 20 GW of capacity in 2032 are announced plans, not capabilities already available.

Those unknowns define the criteria by which the V900 can eventually be judged when commercial systems arrive: how much useful computation can Alibaba actually deliver, at what cost, with what power consumption and for which tasks?

FAQ

Is the Zhenwu V900 already available?

No. Alibaba plans mass production for the first quarter of 2027. The performance figures announced on September 22, 2026 therefore do not yet describe a product broadly deployed with customers.

Is the V900 faster than Nvidia's AI chips?

Public information does not yet support a rigorous answer. Alibaba compares the V900 with its own M890, but the sources reviewed do not provide an independent like-for-like benchmark against Nvidia accelerators using identical workloads, numerical precision and power envelopes.

Why is Alibaba developing chips, models and data centers simultaneously?

Vertical integration can reduce some dependencies, enable closer hardware-software optimization and give Alibaba greater control over costs. In return, the company must master more industrial layers, from silicon and networking to software, energy and data-center operations.