RPG Framework: Guided Self-Improvement Boosts Embodied Agent Success to 95% Without Weight Updates
自主學習免微調!「RPG」框架藉由虛擬練習與自我診斷,將機器人任務成功率提升至 95%
Developing reliable robot skills usually requires extensive human engineering. The RPG (Reconstruct, Practice, Go Real) framework automates this by extracting tasks from offline data, practicing in simulation, diagnosing failures via privileged states, and generating/refining symbolic skills and system prompts without updating neural network weights. RPG boosted task success on 22 manipulation tasks from 28.6% to 95.0%, outperforming GPT-6-based agents, and achieved a 100% success rate in real-world physical tests.
Key points
No Weight Tuning Required
Implements self-improvement solely by optimizing system prompts and symbolic skill libraries, bypassing neural network weight updates.
Autonomous Failure Diagnosis
Diagnoses execution failures during simulation using privileged simulator states and offline videos, automatically creating and refining reusable skills.
Cross-Task Validation
Evaluates candidate changes and merged revisions across multiple tasks before committing them to the reusable library.
Zero-Shot Physical Transfer
Boosted simulation success to 95.0% after 15 rounds and achieved 100% success in 30 physical robot trials after calibration.
How it works
Why it matters
Traditional robotics heavily relies on hand-crafted reward functions and manual tuning. RPG demonstrates that combining LLM reasoning with simulation practice allows robots to autonomously 'learn from failures' without altering underlying model weights. This significantly reduces the development and maintenance costs of deploying embodied AI in real-world environments.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Autonomous skill acquisition: Automatically learning and optimizing complex manipulation skills (e.g., grasping, rotating) in simulation.
- 2Sim-to-Real Transfer: Deploying optimized system prompts and skill libraries directly to physical robot arms with minimal hardware adaptation.
Limitations & caveats
- Strong dependency on simulators: If target tasks are too complex to reconstruct in simulation, the framework cannot perform effective practice.
- Still requires a manual calibration and hardware-adaptation procedure before deploying on physical robots.
Related
DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents
DynaHarness:具備自我演化能力的機器人代理動態實體約束框架
DynaHarness is a dynamic physical framework for self-evolving robots that bridges semantic reasoning and execution through a contract, transforming failures into capability updates.
Skill-Space Shooting: Autonomous Robot Policy Improvement Guided by Foundation Models
技能空間射擊演算法:利用基礎模型引導機器人自主進行策略優化
This research introduces 'skill-space shooting,' a method that leverages foundation models to guide exploration in a reusable skill space, enabling robots to autonomously correct and improve their policies.
AD-WM: Action-Discriminative World Models for Counterfactual MPC
AD-WM:專為反事實預測控制設計的動作辨識世界模型
AD-WM enhances latent world models by learning to distinguish alternative actions from the same state, significantly improving counterfactual control and zero-shot transfer in robotics.