DepthWorld: 3D World Model for Robot Manipulation
DepthWorld:為機器人操控打造的 3D 世界模型
Traditional video-based world models rely only on RGB, producing realistic frames that lack 3D geometric consistency. To solve this, researchers introduced a calibration pipeline to upgrade the DROID dataset into DROID-3D, containing dense metric depth and precise extrinsics. They trained DepthWorld (based on Stable Video Diffusion) using spatial latent tiling to jointly predict multi-view RGB and depth, boosting RGB prediction by +1.48 dB PSNR while providing accurate depth for robotic reasoning.
Key points
DROID-3D Calibrated Dataset
Combines learned stereo depth with a factor graph to recover kinematic parameters, achieving <0.7 px reprojection error on 90% of episodes.
Spatial Latent Tiling
Jointly predicts multi-view RGB and depth via spatial latent tiling without modifying the underlying pretrained VAE.
Depth-Boosted RGB Prediction
Integrating depth supervision improves the quality of RGB predictions by +1.48 dB PSNR under an identical training budget.
How it works
Why it matters
Robotic tasks like planning and policy evaluation depend heavily on true 3D geometry. While prior video models suffer from warping and inconsistency, DepthWorld proves that incorporating depth data not only outputs reliable metric depth for spatial reasoning but also improves RGB generation fidelity, offering a much more grounded simulation environment for Embodied AI.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Robot policy evaluation and multi-step planning
- 2High-fidelity 3D physics and environment simulation
Limitations & caveats
- Relies heavily on custom joint factor graph calibration to generate initial 3D dataset labels.
- Joint multi-view and depth diffusion may introduce increased computational overhead.
Related
QF3: Fast Flow RL with Filtered Q-Gradients
QF3:利用過濾 Q 梯度實現快速流匹配強化學習的機器人控制技術
QF3 is an off-policy RL algorithm that accelerates flow-based policy training with a 10x speedup, enabling humanoid locomotion training from scratch and successful zero-shot hardware transfer.
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
VeriFine:透過協同演化驗證機制實現具身智慧的自我迭代
VeriFine is an agent framework for embodied reasoning that co-evolves the policy and its verification judge through dual loops, overcoming the optimization bottlenecks of static evaluation.
EyeRobot 2.0: Precise Robot Manipulation via Active Gaze Without Wrist Cameras
EyeRobot 2.0:無需手腕相機,用「主動注視」實現精準雙手機器人操控
EyeRobot 2.0 mimics human vision using active gaze with a single stereo camera, enabling precise bimanual manipulation without wrist-mounted cameras.