Huawei's Ascend 960 arrives early in a 4,096-card supernode
Huawei has pulled its Ascend 960 chip forward by three quarters and paired it with the first supernode built on near-packaged optics: 4,096 cards acting as one machine.

China's answer to the rack-scale problem arrived in Shanghai on Wednesday with a timetable Huawei had no business keeping. At Huawei Connect the company said its Ascend 960 chip is running three quarters ahead of schedule and paired it with what it calls the industry's first supernode built on near-packaged optics: a 4,096-card block that Huawei says behaves like one very large computer.
What Huawei actually announced
The headline is the date. Rotating chairman Wang Tao told the keynote that the Ascend 960DT, previously pencilled in for late 2027, will now be ready in the first quarter of 2027, three quarters early, while the higher-spec 960PR moves to the third quarter of 2027, a quarter ahead of plan. Huawei says it will hold to a one-generation-a-year rhythm from there, with the Ascend 970 in 2028 and the 980 in 2029, each roughly doubling compute with more memory bandwidth and capacity alongside it.
The more interesting engineering sits in the wiring. The Ascend 960 supernode is the first of its kind to use near-packaged optics, Huawei's term for parking the optical engine right beside the chip instead of at the edge of the board. Huawei says one supernode holds 4,096 cards and offers up to 8E FP8 of compute and 1PB of HBM memory. Inside it, 5,500 Hi-ONE optical engines do the work of 48,000 800-gigabit modules, which the company says strips out more than 550kW of power draw, doubles the time between failures and takes system availability to 99.8 per cent. Each Hi-ONE engine moves 7.2Tbps, and Huawei claims it is the only mass-produced design of its kind with a built-in light source.
The supporting cast got a refresh too. The Kunpeng supernode now scales to 4,096 nodes with a 256TB unified memory pool, which Huawei says starts 100,000-class agent sandboxes 30 times faster than conventional servers, and OceanStor M900 adds a petabyte-scale key-value cache tier that the company claims stretches SSD write life 16-fold. Huawei's stated ceiling for a single cluster is 512,000 cards across a two-layer fabric, rising to a million Ascend supernodes with its multi-rail topology. On software, it says Ascend is now an installable platform on PyTorch's own site and that openEuler has passed 20 million installations.
What is a supernode, in plain English?
A normal AI rack is a pile of servers chattering to each other over a network, and the network is where the time disappears. Huawei's own simulation work reckons communication eats more than 40 per cent of training time on a 100,000-card cluster built from eight-card servers, which drags down how much of the silicon is actually doing maths. A supernode flattens that topology: many compute boards are wired together so tightly that they share one pool of memory and present themselves as a single machine. Swap the eight-card servers for 4,000-card supernodes and Huawei reckons that utilisation figure improves 2.75 times. It is the same argument Nvidia makes with its own racks, and the same reason the interconnects, rather than the chips, have become the real battleground.
What Huawei did not say
Almost every number above is Huawei's own, with no independent benchmark, no price, no yield figure and no named customer attached, so treat the 2.75x and the 99.8 per cent as claims rather than measurements. There is a standards wrinkle as well: Huawei says it has proposed a near-packaged optics standard project at the Optical Internetworking Forum, which is a useful reminder that the ecosystem around this technology is still being written, and that shipping early is not the same as shipping in volume. This is a roadmap and a specification sheet; the deployment receipts come later.
Our opinion
This is the most concrete thing Huawei has shown on AI infrastructure in a while, and the timetable is the story: you do not pull silicon forward by three quarters unless the pressure is genuinely on. But “ready” in a keynote and “available” on a purchase order are different words, and until there are independent benchmarks, published prices and a customer willing to put its name on 4,096 cards, this remains a promise wearing a product's clothes. Worth watching closely, with the scepticism still set for a Q1 2027 that has not happened yet.
- Huawei Connect 2026 opened in Shanghai on 17 September, with rotating chairman Wang Tao giving the keynote
- Huawei says the Ascend 960DT is now due in Q1 2027, three quarters earlier than planned, with the 960PR following in Q3 2027
- It promises a new Ascend generation every year after that: the 970 in 2028 and the 980 in 2029
- The Ascend 960 supernode is the first to use near-packaged optics, at 4,096 cards, up to 8E FP8 of compute and 1PB of HBM memory
- Huawei claims 5,500 Hi-ONE optical engines replace 48,000 800-gigabit modules, cutting over 550kW and lifting availability to 99.8%
- Its stated ceiling for one cluster is 512,000 cards, or a million Ascend supernodes with multi-rail topology
- Every performance, power and reliability figure is Huawei's own; no independent benchmark, price or named customer was published