A Balanced Data Diet: Success Guided Sampling for Mega-Scale Robot RL
超大規模機器人強化學習的平衡數據飲食:成功引導採樣 SGS
Traditional mega-scale parallel RL relies on uniform simulator resets, wasting compute on task configurations that are either already mastered or impossible to attempt. The authors introduce Success Guided Sampling (SGS), an adaptive sampling algorithm that dynamically directs simulation resets toward the frontier of the policy's current capabilities. Evaluated in up to 2^20 (over 1 million) parallel environments, SGS solves complex multi-terrain locomotion and contact-rich assembly tasks, with policies distilled into RGB vision models for zero-shot real-robot deployment.
Key points
Exploration Bottleneck in Parallel RL
Naively scaling parallel environments with uniform resets wastes batch experience on trivial or impossible tasks.
Success Guided Sampling
SGS acts as an adaptive sampler that concentrates training on task configurations along the frontier of policy capabilities.
Scaled up to 1 Million Environments
Proven effective in up to 2^20 parallel environments, solving multi-terrain quadruped locomotion and contact-rich assembly.
Zero-Shot Real-Robot Deployment
Learned manipulation policies are distilled into RGB-based vision models, enabling zero-shot transfer onto real robot hardware.
How it works
Why it matters
As simulation engines reach the million-environment scale, compute efficiency becomes the primary bottleneck. SGS reduces the need for heavy hand-engineered reward shaping and human demonstrations. By adaptively sampling task frontiers, SGS unlocks true scaling for reinforcement learning, advancing general-purpose locomotion and manipulation on real hardware.
Who it affects
- AI Researcher
- AI Developer
- Student & Learner
How to use it
- 1Multi-terrain agile locomotion for quadrupedal robots
- 2Contact-rich precision assembly for robotic manipulators
- 3Adaptive training resource allocation in mega-scale parallel physics simulations
Limitations & caveats
- Requires simulation infrastructure supporting dynamic resets and real-time capability frontier tracking.
- Real-world success depends heavily on the quality of distillation into RGB-based vision policies across the sim-to-real gap.
Related

5 Steps to Build SimReady Robotics Assets with Frontier AI Models
5 步驟建構 SimReady 機器人資產:結合前沿 AI 模型與 NVIDIA Omniverse
A structured 5-step workflow leveraging frontier AI agents to convert, configure, and validate CAD assets into physics-ready SimReady robotics models inside NVIDIA Isaac Sim.
CSF: Contextual Safety Filtering for Text-Conditioned Motion Generators
CSF:結合場景語境的機器人動作生成安全過濾技術
CSF is a training-free framework that grounds natural-language safety rules in scene context using Control Barrier Functions to prevent unsafe robot motions.

The Machines That Make the Machines: How NVIDIA Automates GB300 Tester Tray Assembly
機器造機器:NVIDIA 如何用 AI 與實體控制自動組裝 GB300 測試托盤
NVIDIA Seattle Robotics Lab shares insights from automating GB300 superchip tester tray assembly, highlighting the synergy between classical control engineering, smart mechanical design, and RL.