Aivora
arXivRoboticsIntermediate

DepthWorld: 3D World Model for Robot Manipulation

DepthWorld:為機器人操控打造的 3D 世界模型

2 min read
DepthWorld: 3D World Model for Robot Manipulation
The 30-second version

Traditional video-based world models rely only on RGB, producing realistic frames that lack 3D geometric consistency. To solve this, researchers introduced a calibration pipeline to upgrade the DROID dataset into DROID-3D, containing dense metric depth and precise extrinsics. They trained DepthWorld (based on Stable Video Diffusion) using spatial latent tiling to jointly predict multi-view RGB and depth, boosting RGB prediction by +1.48 dB PSNR while providing accurate depth for robotic reasoning.

Key points

01

DROID-3D Calibrated Dataset

Combines learned stereo depth with a factor graph to recover kinematic parameters, achieving <0.7 px reprojection error on 90% of episodes.

02

Spatial Latent Tiling

Jointly predicts multi-view RGB and depth via spatial latent tiling without modifying the underlying pretrained VAE.

03

Depth-Boosted RGB Prediction

Integrating depth supervision improves the quality of RGB predictions by +1.48 dB PSNR under an identical training budget.

How it works

DepthWorld Training and Inference Workflow
Input pipelineYields 3D labelsJoint latent trainingPredicts outputRaw DROID VideosFactor GraphCalibrationDROID-3D Metric DepthDepthWorld DiffusionModelConsistent Multi-viewRGB & Depth

Why it matters

Robotic tasks like planning and policy evaluation depend heavily on true 3D geometry. While prior video models suffer from warping and inconsistency, DepthWorld proves that incorporating depth data not only outputs reliable metric depth for spatial reasoning but also improves RGB generation fidelity, offering a much more grounded simulation environment for Embodied AI.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Robot policy evaluation and multi-step planning
  2. 2High-fidelity 3D physics and environment simulation

Limitations & caveats

  • Relies heavily on custom joint factor graph calibration to generate initial 3D dataset labels.
  • Joint multi-view and depth diffusion may introduce increased computational overhead.

Related

QF3: Fast Flow RL with Filtered Q-Gradients
arXivRobotics

QF3: Fast Flow RL with Filtered Q-Gradients

QF3:利用過濾 Q 梯度實現快速流匹配強化學習的機器人控制技術

QF3 is an off-policy RL algorithm that accelerates flow-based policy training with a 10x speedup, enabling humanoid locomotion training from scratch and successful zero-shot hardware transfer.

2 min read
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
arXivRobotics

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine:透過協同演化驗證機制實現具身智慧的自我迭代

VeriFine is an agent framework for embodied reasoning that co-evolves the policy and its verification judge through dual loops, overcoming the optimization bottlenecks of static evaluation.

2 min read