Trending: On-device modelsSearch
iHeartGeek
iTECH

NVIDIA DSX MaxLPS lifts AI factory throughput 49.2% on the same power budget

A joint NVIDIA and Nscale evaluation of GB300 NVL72 racks found DSX MaxLPS added 52 GPUs inside the same 264.4 kW budget, raising normalised throughput 49.2% but pushing 99th percentile time to first token up 17%.

Render of an AI data centre connected by glowing green power lines to a dial, a battery unit and a transmission tower

NVIDIA says its DSX MaxLPS power-sharing layer can squeeze up to 40% more GPUs into an AI factory's existing approved power budget, and it has published the measurements from a joint trial with Nscale to back that up.

What's actually going on

AI data centres are provisioned for the unlikely moment when every accelerator draws peak power at once. That buffer is sensible, but AI workloads rarely behave that way: training runs cycle through compute, communication and checkpointing, while inference alternates between prefill, decode, memory-bound work and idle. The result is reserved capacity that one node cannot lend to another even when the facility as a whole sits below its limit.

DSX MaxLPS monitors real consumption and dynamically reallocates power across participating resources while holding the operator's aggregate budget and policy boundaries. NVIDIA is careful to frame this as coordinated allocation inside the same managed power envelope rather than an increase in the site's supply. The control layer comes from the company's Dynamic Power Software.

The evaluation ran at Nscale's Verne campus in Keflavík, Iceland, on renewable power. The workload was Kimi K2.5 in FP4 on Blackwell Ultra GB300 NVL72 systems, using NVIDIA Dynamo and TensorRT LLM, with an 8K input sequence length and a 1K output. Jobs were spread across four racks, with each distributed workload confined to a single rack so cross-rack differences could not flatter the result.

The numbers

The static baseline used 35 four-GPU nodes, or 140 GPUs, running two high-throughput instances of 52 GPUs each and one low-latency instance of 36 GPUs. The DSX MaxLPS configuration used 48 four-GPU nodes, or 192 GPUs, adding a third 52-GPU high-throughput instance while keeping the same 36-GPU low-latency one. Both ran inside the same 264.4 kW provisioned power budget.

Normalised aggregate throughput rose 49.2%. Throughput per provisioned watt climbed from 4.10 to 6.12 tokens per second per watt. Per-instance throughput for the high-throughput and low-latency workloads was effectively unchanged, which is the part that matters for service quality.

Latency tells the more cautious half of the story. Median and 75th percentile latency stayed within 5% of the baseline, but 99th percentile time to first token increased 17%. That is a real trade-off, and NVIDIA is explicit that operators need to watch tail latency alongside the capacity gain rather than reading the throughput figure on its own.

For rollout, NVIDIA recommends a five-stage process: define the managed boundary, establish a representative baseline, introduce policies conservatively, add capacity incrementally with testing at every stage, and only set production operating limits once every objective has been met.

Our opinion

Power is the binding constraint on AI build-outs now, not silicon supply, so an approach that buys nearly half again in throughput without asking the grid for more megawatts is the most useful kind of engineering. The detail that makes this credible is the 17% rise in 99th percentile time to first token. Vendors rarely volunteer their worst-case regression, and putting it in the same table as a 49.2% gain is the difference between a press release and a benchmark. Operators should treat the headline as a capacity ceiling to validate rather than a setting to switch on, but the direction of travel is clear: the next generation of AI factory efficiency will come from scheduling power rather than building more of it.