Aivora
NVIDIA DeveloperAI HardwareIntermediate

Unlocking Stranded Capacity: How NVIDIA DSX MaxLPS Boosts AI Factory Throughput by 49%

破解電力瓶頸:NVIDIA DSX MaxLPS 如何釋放 AI 工廠高達 40% 的隱形算力

2 min read
Unlocking Stranded Capacity: How NVIDIA DSX MaxLPS Boosts AI Factory Throughput by 49%
The 30-second version

Traditional AI data centers statically reserve power based on peak GPU consumption, leaving valuable capacity stranded. NVIDIA's DSX MaxLPS solves this with dynamic, policy-governed power allocation. Tested at Nscale's Iceland facility with Kimi K2.5 workloads on Blackwell-based GB300 NVL72 systems, MaxLPS enabled a 37% increase in active GPUs (from 140 to 192) within a fixed 264.4 kW budget, achieving a 49.2% surge in aggregate token throughput without sacrificing average application performance.

Key points

01

Dynamic Power Allocation

Monitors real-time consumption across resources and reallocates unused power dynamically instead of relying on rigid static peak reservations.

02

Massive Throughput Gain

Delivers a 49.2% increase in both aggregate throughput and tokens-per-watt without raising the pre-approved facility power budget.

03

Stable Core Performance

Median and P75 latencies remained within 5% of the baseline, preserving standard service quality across running instances.

04

Structured Stage Validation

Recommends a staged validation process—mapping topologies, establishing baselines, and gradually adding capacity—to guarantee safety under stress.

How it works

Static Baseline vs. DSX MaxLPS Performance Comparison
Static baseline (靜態基準)DSX MaxLPS (動態調配)
Managed GPUs140192 (+37.1%)
Aggregate throughput1,084,5031,618,443 (+49.2%)
Throughput per watt4.10 tokens/s/W6.12 tokens/s/W (+49.2%)
Power budget utilization62.9%75.2% (+12.3 pp)
Total measured power166.2 kW198.9 kW (+19.7%)

Why it matters

As AI models scale exponentially, power availability has become the ultimate bottleneck for data center expansion. DSX MaxLPS proves that software-defined hardware management can reclaim massive amount of stranded power. By allowing operators to pack up to 40% more GPUs into existing envelopes, it drastically reduces infrastructure costs per token and bypasses utility constraints without waiting for grid upgrades.

Who it affects

  • Enterprise Leader
  • AI Developer
  • Startup Founder

How to use it

  1. 1Optimizing AI factories facing strict utility grid power limits
  2. 2Managing heterogeneous fleet power profiles mixing training and inference workloads
  3. 3Sustainable, high-efficiency LLM deployment in eco-friendly data centers

Limitations & caveats

  • Tail latency (P99 time to first token) increased by 17%, which demands careful evaluation for ultra-low latency services.
  • Highly dependent on precise, real-time telemetry; delayed or missing data can undermine fleet-level control loop safety.
  • Reclaiming headroom increases overall power draw (+19.7%), meaning physical cooling and distribution must still support peak capacity lifecycle targets.

Related