NVIDIA Expands AI Storage Acceleration with cuObject and SCADA Server SDK
NVIDIA 推出 cuObject 與 SCADA 伺服器 SDK:實現免經 CPU 的 GPU 直連快取與物件儲存加速

Traditional storage access paths that copy data through the server's CPU memory limit AI training and inference speeds. NVIDIA is addressing this by expanding xio-sig to include cuObject (for object storage) and introducing the SCADA Server SDK (for fine-grained, GPU-initiated storage requests). These open-standard solutions enable GPUs to access storage directly over RDMA without CPU overhead. Tech leaders like Google Cloud, Microsoft, and IBM are backing these initiatives to build an interoperable high-performance storage ecosystem.
Key points
cuObject GA & Expansion
NVIDIA is expanding xio-sig to include cuObject alongside cuFile, standardizing accelerated object storage access via RDMA.
SCADA Server SDK
Enables storage providers to build servers responding to high-throughput, fine-grained, GPU-initiated requests over RDMA.
Zero-Copy RDMA Bypass
Leverages DPU/NIC-accelerated RDMA to bypass CPU memory copies, delivering higher throughput and lower latency for AI workloads.
Open Industry Standard
Supported by Google Cloud, Microsoft, and IBM through the Storage-Next initiative to build an interoperable GPU storage standard.
How it works
Why it matters
As LLMs, RAG, and large-scale AI training demand massive data throughput, CPU-bound storage access becomes a critical bottleneck. Standardizing cuObject and SCADA under open initiatives allows diverse storage vendors to deliver unified, ultra-low-latency RDMA storage directly to GPUs. This optimizes large-scale systems for semantic search, recommendation engines, and real-time fraud detection.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Accelerating data loading during large-scale AI model training and fine-tuning.
- 2Handling fine-grained query lookups for semantic search and RAG vector databases.
- 3High-throughput, low-latency data fetching for recommender systems and fraud detection.
Limitations & caveats
- Requires RDMA-capable hardware such as NVIDIA ConnectX NICs or BlueField DPUs, increasing initial setup costs.
- Some governance frameworks and core kernel code are still undergoing compliance testing and board review.
Related

Unlocking Stranded Capacity: How NVIDIA DSX MaxLPS Boosts AI Factory Throughput by 49%
破解電力瓶頸:NVIDIA DSX MaxLPS 如何釋放 AI 工廠高達 40% 的隱形算力
A joint NVIDIA and Nscale evaluation demonstrates that DSX MaxLPS dynamically allocates power to deploy 37% more GPUs and boost LLM throughput by 49% within the same power budget.

NVIDIA Launches Open-Source NVCRE: Automating GPU Cluster Readiness Validation Before AI Workloads Run
NVIDIA 開源 NVCRE:在 AI 工作負載上線前,自動驗證 GPU 叢集準備狀態
Traditional GPU health checks often miss performance bottlenecks in distributed training. NVIDIA's open-source NVCRE is a Kubernetes controller that runs active, topology-aware workloads to pinpoint failing nodes before production starts.

Overcoming Confidential Computing Overheads: Optimizing Private LLM Inference on NVIDIA Blackwell
突破機密運算效能瓶頸:NVIDIA Blackwell 與 TensorRT-LLM 的隱私推理優化
This article explains how NVIDIA optimizes TensorRT-LLM on Blackwell GPUs to mitigate Confidential Computing overheads, retaining up to 98.2% of throughput for DeepSeek-R1 private inference.