EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
EmbodiedRSI:首個透過「假設引導」實現自主演化與持續學習的機器人框架
Robot foundation models often degrade when task instructions or environments change, while retraining via teleoperation is highly expensive. EmbodiedRSI overcomes this with a Fast-Slow Dual-System. It maintains competing code and skill hypotheses within a Hypothesis Graph, runs physical trials optimized by Value-of-Information (VoI), and performs Code-Skill Co-Evolution. Guided by Hierarchical Memory, it achieves 77% success on RoboCasa365 and zero-shot transfers to real robots with 71.3% success.
Key points
Fast-Slow Dual-System
Combines a fast execution system with a slow learning system that maintains hierarchical memories and refines agent behaviors over time.
Hypothesis Graph & VoI Selection
Maintains a graph of competing hypotheses and uses Value-of-Information (VoI) to select physical experiments that resolve system uncertainties.
Code-Skill Co-Evolution
Simultaneously refines high-level control code and low-level skill parameters based on real-world interaction outcomes without human intervention.
Strong Cross-Domain Transfer
Transfers zero-shot from simulation to physical robots, achieving a 71.3% overall success rate across challenging real-world manipulation tasks.
How it works
Why it matters
Data collection for Embodied AI is historically expensive. EmbodiedRSI demonstrates a closed-loop scientific paradigm where robots propose hypotheses, design physical experiments, and self-correct. By drastically reducing reliance on human demonstrations and teleoperation, this research paves the way for self-evolving, lifelong learning robots in highly dynamic, real-world environments.
Who it affects
- AI Researcher
- AI Developer
- Enterprise Leader
How to use it
- 1Kitchen and service robots autonomously exploring and learning to manipulate novel utensils and unseen object layouts.
- 2Industrial robotic arms performing self-debugging and optimizing precision skills for newly introduced parts without down-time.
Limitations & caveats
- During early exploration phases, physical execution based on unverified hypotheses may carry collision and safety risks.
- Scaling to massive, multi-task settings may increase the computational overhead of maintaining the Hypothesis Graph and calculating VoI.
Related

The Machines That Make the Machines: How NVIDIA Automates GB300 Tester Tray Assembly
機器造機器:NVIDIA 如何用 AI 與實體控制自動組裝 GB300 測試托盤
NVIDIA Seattle Robotics Lab shares insights from automating GB300 superchip tester tray assembly, highlighting the synergy between classical control engineering, smart mechanical design, and RL.
Long-WAM: Scaling Context for Real-Time World-Action Models
Long-WAM:突破即時控制延遲,擴展機器人世界動作模型的長上下文
Long-WAM scales the visual history context of world-action models under real-time constraints, demonstrating that autoregressive pretraining is key to unlocking the power of long video memories.
"Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models
機器人控制的「文字敏感症」:為何一個詞能讓 VLA 模型成功率從 100% 跌到 2%?
This study exposes the extreme sensitivity of Vision-Language-Action (VLA) models to instruction phrasing and introduces a zero-shot framework that uses LLMs to distill rewriting rules, significantly boosting robotic task success without retraining.