Aivora
arXivRoboticsAdvanced

EyeRobot 2.0: Precise Robot Manipulation via Active Gaze Without Wrist Cameras

EyeRobot 2.0:無需手腕相機,用「主動注視」實現精準雙手機器人操控

2 min read
EyeRobot 2.0: Precise Robot Manipulation via Active Gaze Without Wrist Cameras
The 30-second version

Traditional robot manipulation relies heavily on wrist cameras, which add hardware overhead and fail during occlusions. EyeRobot 2.0 introduces Active Visual Fixation (AVF) using a single stereo camera that swivels to physically center its gaze on 3D points, processing images foveally by focusing visual tokens on the center. Trained hierarchically via RL and BC, EyeRobot 2.0 outperforms passive stereo by 40% in real-world trials and even beats wrist-camera systems when grasped objects cause visual occlusions.

Key points

01

Active Visual Fixation

Physically swivels a stereo camera to center on a 3D fixation point, allocating more visual tokens to the center to mimic foveal vision.

02

Hierarchical Control

A low-level gaze servoing policy guides focal adjustment, while a high-level target selector co-trains with the gripper policy to plan fixation paths.

03

Fixation-Relative Canonicalization

Transforms gripper information into a fixation-relative SE(3) frame, compacting the action distribution for easier learning.

04

Robust to Occlusions

When held objects block wrist-camera views, EyeRobot 2.0 achieves 48% success, more than doubling the 22% of wrist-camera baselines.

How it works

EyeRobot 2.0 System Architecture and Information Flow
Task progressFixation targetSwivel commandsRaw imagesFoveal featuresSE(3) ActionsEnv & ObjectsTarget SelectorGaze ServoingStereo CameraFoveal ProcessingGripper Policy

Why it matters

This research proves that active gaze can successfully replace hardware-heavy wrist cameras. In complex bimanual tasks, wrist-mounted cameras are highly prone to occlusion by the robot's own arms or grasped objects. EyeRobot 2.0 reduces hardware complexity and cost while offering a highly robust, occlusion-free alternative for robotic manipulation by mimicking human eye movements.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Fine-grained bimanual assembly (e.g., electronic component insertion) where wrist cameras are easily occluded.
  2. 2Low-cost collaborative robot vision setups that seek to minimize hardware complexity.

Limitations & caveats

  • Requires precise control over physical camera swiveling mechanisms, adding mechanical complexity to the robot's head.
  • Relies heavily on joint RL and BC training, meaning its generalization to completely unseen objects requires further evaluation.

Related

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents
arXivRobotics

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

DynaHarness:具備自我演化能力的機器人代理動態實體約束框架

DynaHarness is a dynamic physical framework for self-evolving robots that bridges semantic reasoning and execution through a contract, transforming failures into capability updates.

2 min read