Aivora
NVIDIA DeveloperAI HardwareAdvanced

NVIDIA Expands AI Storage Acceleration with cuObject and SCADA Server SDK

NVIDIA 推出 cuObject 與 SCADA 伺服器 SDK:實現免經 CPU 的 GPU 直連快取與物件儲存加速

2 min read
NVIDIA Expands AI Storage Acceleration with cuObject and SCADA Server SDK
The 30-second version

Traditional storage access paths that copy data through the server's CPU memory limit AI training and inference speeds. NVIDIA is addressing this by expanding xio-sig to include cuObject (for object storage) and introducing the SCADA Server SDK (for fine-grained, GPU-initiated storage requests). These open-standard solutions enable GPUs to access storage directly over RDMA without CPU overhead. Tech leaders like Google Cloud, Microsoft, and IBM are backing these initiatives to build an interoperable high-performance storage ecosystem.

Key points

01

cuObject GA & Expansion

NVIDIA is expanding xio-sig to include cuObject alongside cuFile, standardizing accelerated object storage access via RDMA.

02

SCADA Server SDK

Enables storage providers to build servers responding to high-throughput, fine-grained, GPU-initiated requests over RDMA.

03

Zero-Copy RDMA Bypass

Leverages DPU/NIC-accelerated RDMA to bypass CPU memory copies, delivering higher throughput and lower latency for AI workloads.

04

Open Industry Standard

Supported by Google Cloud, Microsoft, and IBM through the Storage-Next initiative to build an interoperable GPU storage standard.

How it works

GPU-Initiated Direct Storage Architecture (SCADA/cuObject)
Initiates RequestBypass CPU CopyData AccessDirect to GPUHost CPU (Bypassed)NIC / DPU RDMASCADA ServerFile/Object StorageGPU Memory

Why it matters

As LLMs, RAG, and large-scale AI training demand massive data throughput, CPU-bound storage access becomes a critical bottleneck. Standardizing cuObject and SCADA under open initiatives allows diverse storage vendors to deliver unified, ultra-low-latency RDMA storage directly to GPUs. This optimizes large-scale systems for semantic search, recommendation engines, and real-time fraud detection.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Accelerating data loading during large-scale AI model training and fine-tuning.
  2. 2Handling fine-grained query lookups for semantic search and RAG vector databases.
  3. 3High-throughput, low-latency data fetching for recommender systems and fraud detection.

Limitations & caveats

  • Requires RDMA-capable hardware such as NVIDIA ConnectX NICs or BlueField DPUs, increasing initial setup costs.
  • Some governance frameworks and core kernel code are still undergoing compliance testing and board review.

Related