Unlocking Stranded Capacity: How NVIDIA DSX MaxLPS Boosts AI Factory Throughput by 49%
破解電力瓶頸:NVIDIA DSX MaxLPS 如何釋放 AI 工廠高達 40% 的隱形算力

Traditional AI data centers statically reserve power based on peak GPU consumption, leaving valuable capacity stranded. NVIDIA's DSX MaxLPS solves this with dynamic, policy-governed power allocation. Tested at Nscale's Iceland facility with Kimi K2.5 workloads on Blackwell-based GB300 NVL72 systems, MaxLPS enabled a 37% increase in active GPUs (from 140 to 192) within a fixed 264.4 kW budget, achieving a 49.2% surge in aggregate token throughput without sacrificing average application performance.
Key points
Dynamic Power Allocation
Monitors real-time consumption across resources and reallocates unused power dynamically instead of relying on rigid static peak reservations.
Massive Throughput Gain
Delivers a 49.2% increase in both aggregate throughput and tokens-per-watt without raising the pre-approved facility power budget.
Stable Core Performance
Median and P75 latencies remained within 5% of the baseline, preserving standard service quality across running instances.
Structured Stage Validation
Recommends a staged validation process—mapping topologies, establishing baselines, and gradually adding capacity—to guarantee safety under stress.
How it works
| Static baseline (靜態基準) | DSX MaxLPS (動態調配) | |
|---|---|---|
| Managed GPUs | 140 | 192 (+37.1%) |
| Aggregate throughput | 1,084,503 | 1,618,443 (+49.2%) |
| Throughput per watt | 4.10 tokens/s/W | 6.12 tokens/s/W (+49.2%) |
| Power budget utilization | 62.9% | 75.2% (+12.3 pp) |
| Total measured power | 166.2 kW | 198.9 kW (+19.7%) |
Why it matters
As AI models scale exponentially, power availability has become the ultimate bottleneck for data center expansion. DSX MaxLPS proves that software-defined hardware management can reclaim massive amount of stranded power. By allowing operators to pack up to 40% more GPUs into existing envelopes, it drastically reduces infrastructure costs per token and bypasses utility constraints without waiting for grid upgrades.
Who it affects
- Enterprise Leader
- AI Developer
- Startup Founder
How to use it
- 1Optimizing AI factories facing strict utility grid power limits
- 2Managing heterogeneous fleet power profiles mixing training and inference workloads
- 3Sustainable, high-efficiency LLM deployment in eco-friendly data centers
Limitations & caveats
- Tail latency (P99 time to first token) increased by 17%, which demands careful evaluation for ultra-low latency services.
- Highly dependent on precise, real-time telemetry; delayed or missing data can undermine fleet-level control loop safety.
- Reclaiming headroom increases overall power draw (+19.7%), meaning physical cooling and distribution must still support peak capacity lifecycle targets.
Related

NVIDIA Launches Open-Source NVCRE: Automating GPU Cluster Readiness Validation Before AI Workloads Run
NVIDIA 開源 NVCRE:在 AI 工作負載上線前,自動驗證 GPU 叢集準備狀態
Traditional GPU health checks often miss performance bottlenecks in distributed training. NVIDIA's open-source NVCRE is a Kubernetes controller that runs active, topology-aware workloads to pinpoint failing nodes before production starts.

Overcoming Confidential Computing Overheads: Optimizing Private LLM Inference on NVIDIA Blackwell
突破機密運算效能瓶頸:NVIDIA Blackwell 與 TensorRT-LLM 的隱私推理優化
This article explains how NVIDIA optimizes TensorRT-LLM on Blackwell GPUs to mitigate Confidential Computing overheads, retaining up to 98.2% of throughput for DeepSeek-R1 private inference.