EyeRobot 2.0: Precise Robot Manipulation via Active Gaze Without Wrist Cameras
EyeRobot 2.0:無需手腕相機,用「主動注視」實現精準雙手機器人操控
Traditional robot manipulation relies heavily on wrist cameras, which add hardware overhead and fail during occlusions. EyeRobot 2.0 introduces Active Visual Fixation (AVF) using a single stereo camera that swivels to physically center its gaze on 3D points, processing images foveally by focusing visual tokens on the center. Trained hierarchically via RL and BC, EyeRobot 2.0 outperforms passive stereo by 40% in real-world trials and even beats wrist-camera systems when grasped objects cause visual occlusions.
Key points
Active Visual Fixation
Physically swivels a stereo camera to center on a 3D fixation point, allocating more visual tokens to the center to mimic foveal vision.
Hierarchical Control
A low-level gaze servoing policy guides focal adjustment, while a high-level target selector co-trains with the gripper policy to plan fixation paths.
Fixation-Relative Canonicalization
Transforms gripper information into a fixation-relative SE(3) frame, compacting the action distribution for easier learning.
Robust to Occlusions
When held objects block wrist-camera views, EyeRobot 2.0 achieves 48% success, more than doubling the 22% of wrist-camera baselines.
How it works
Why it matters
This research proves that active gaze can successfully replace hardware-heavy wrist cameras. In complex bimanual tasks, wrist-mounted cameras are highly prone to occlusion by the robot's own arms or grasped objects. EyeRobot 2.0 reduces hardware complexity and cost while offering a highly robust, occlusion-free alternative for robotic manipulation by mimicking human eye movements.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Fine-grained bimanual assembly (e.g., electronic component insertion) where wrist cameras are easily occluded.
- 2Low-cost collaborative robot vision setups that seek to minimize hardware complexity.
Limitations & caveats
- Requires precise control over physical camera swiveling mechanisms, adding mechanical complexity to the robot's head.
- Relies heavily on joint RL and BC training, meaning its generalization to completely unseen objects requires further evaluation.
Related
RPG Framework: Guided Self-Improvement Boosts Embodied Agent Success to 95% Without Weight Updates
自主學習免微調!「RPG」框架藉由虛擬練習與自我診斷,將機器人任務成功率提升至 95%
The RPG framework enables autonomous robot improvement without updating model weights, boosting task success from 28.6% to 95.0% through simulation practice, failure diagnosis, and real-world transfer.
DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents
DynaHarness:具備自我演化能力的機器人代理動態實體約束框架
DynaHarness is a dynamic physical framework for self-evolving robots that bridges semantic reasoning and execution through a contract, transforming failures into capability updates.
Skill-Space Shooting: Autonomous Robot Policy Improvement Guided by Foundation Models
技能空間射擊演算法:利用基礎模型引導機器人自主進行策略優化
This research introduces 'skill-space shooting,' a method that leverages foundation models to guide exploration in a reusable skill space, enabling robots to autonomously correct and improve their policies.