Aivora
arXivRoboticsAdvanced

EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

EmbodiedRSI:首個透過「假設引導」實現自主演化與持續學習的機器人框架

2 min read
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
The 30-second version

Robot foundation models often degrade when task instructions or environments change, while retraining via teleoperation is highly expensive. EmbodiedRSI overcomes this with a Fast-Slow Dual-System. It maintains competing code and skill hypotheses within a Hypothesis Graph, runs physical trials optimized by Value-of-Information (VoI), and performs Code-Skill Co-Evolution. Guided by Hierarchical Memory, it achieves 77% success on RoboCasa365 and zero-shot transfers to real robots with 71.3% success.

Key points

01

Fast-Slow Dual-System

Combines a fast execution system with a slow learning system that maintains hierarchical memories and refines agent behaviors over time.

02

Hypothesis Graph & VoI Selection

Maintains a graph of competing hypotheses and uses Value-of-Information (VoI) to select physical experiments that resolve system uncertainties.

03

Code-Skill Co-Evolution

Simultaneously refines high-level control code and low-level skill parameters based on real-world interaction outcomes without human intervention.

04

Strong Cross-Domain Transfer

Transfers zero-shot from simulation to physical robots, achieving a 71.3% overall success rate across challenging real-world manipulation tasks.

How it works

EmbodiedRSI Self-Evolving Workflow Architecture
Provide candidate hypothesesRun diagnostic trialsFeed back physical outcomesRefine hypothesesConsolidate to memoryGuide future reasoningVoI ExperimentSelectionPhysical ExplorationCode-Skill Co-EvolutionHypothesis GraphHierarchical Memory

Why it matters

Data collection for Embodied AI is historically expensive. EmbodiedRSI demonstrates a closed-loop scientific paradigm where robots propose hypotheses, design physical experiments, and self-correct. By drastically reducing reliance on human demonstrations and teleoperation, this research paves the way for self-evolving, lifelong learning robots in highly dynamic, real-world environments.

Who it affects

  • AI Researcher
  • AI Developer
  • Enterprise Leader

How to use it

  1. 1Kitchen and service robots autonomously exploring and learning to manipulate novel utensils and unseen object layouts.
  2. 2Industrial robotic arms performing self-debugging and optimizing precision skills for newly introduced parts without down-time.

Limitations & caveats

  • During early exploration phases, physical execution based on unverified hypotheses may carry collision and safety risks.
  • Scaling to massive, multi-task settings may increase the computational overhead of maintaining the Hypothesis Graph and calculating VoI.

Related

Long-WAM: Scaling Context for Real-Time World-Action Models
arXivRobotics

Long-WAM: Scaling Context for Real-Time World-Action Models

Long-WAM:突破即時控制延遲,擴展機器人世界動作模型的長上下文

Long-WAM scales the visual history context of world-action models under real-time constraints, demonstrating that autoregressive pretraining is key to unlocking the power of long video memories.

2 min read