Aivora
arXivRoboticsAdvanced

Long-WAM: Scaling Context for Real-Time World-Action Models

Long-WAM:突破即時控制延遲,擴展機器人世界動作模型的長上下文

2 min read
Long-WAM: Scaling Context for Real-Time World-Action Models
The 30-second version

This paper presents Long-WAM, a framework designed to scale visual history context for robotic world-action models under strict real-time constraints. The authors reveal that longer history only benefits control performance when the video foundation is pretrained autoregressively (AR). By first learning causal prediction from unlabeled videos and preserving this structure during action fine-tuning, combined with streaming encoding and asynchronous execution, Long-WAM achieves a low inference latency of 107.4 ms on RTX 5090. It enables highly dynamic manipulation, such as dynamic cup stacking, on physical platforms.

Key points

01

Autoregressive pretraining unlocks history

Simply scaling history context yields no net gain unless the video foundation is pretrained autoregressively, establishing a causal history-to-future relationship.

02

Longer context yields higher success

Increasing context from 0.0 to 19.2 seconds raises the success rate from 63.3% to 78.7% on RoboCasa GR-1, proving the power of temporal memory.

03

System optimization for real-time control

Through streaming observation encoding and asynchronous execution, Long-WAM runs at 107.4 ms per action chunk on an RTX 5090 without dropping future predictions.

04

Physical breakthrough in dynamic tasks

Deployed on Unitree G1 and YAM robots, achieving 95% success on dynamic cup stacking, whereas baseline methods failed completely.

How it works

Long-WAM Two-Stage Development and Real-Time Execution Pipeline
Input videosTransfer causal memoryDeploy to hardwareOutput actionsUnlabeled Videos (Robot& Egocentric)Stage 1: AR Pretraining(Learn Futur…Stage 2: WAM Adaptation(Introduce Actio…Real-Time Deployment(Streaming + Asyn…Physical Control (G1)(19.2s Memory / 107ms)

Why it matters

Standard robotic policies often sacrifice visual history to maintain low latency, rendering them fragile in dynamic or long-horizon tasks. Long-WAM addresses this by demonstrating that autoregressive pretraining is essential to utilize long temporal context. By combining this insight with system-level streaming optimization, it allows robots to maintain a long memory without latency penalties, paving the way for humanoid robots to handle highly dynamic, multi-step real-world workloads.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Dynamic and rapid manipulation tasks on physical humanoid robots (e.g., dynamic cup stacking with Unitree G1).
  2. 2Long-horizon and multi-step composite manipulation tasks requiring consistent temporal memory.
  3. 3Memory-informed low-level executor complementing high-level planners in complex environments.

Limitations & caveats

  • Heavy reliance on high-performance hardware to sustain sub-110ms streaming inference for action and video predictions.
  • Requires large volumes of unlabeled robot and egocentric video data for the critical autoregressive pretraining phase.

Related

RoboJEPA: Scaling Robotic Latent World Models
arXivRobotics

RoboJEPA: Scaling Robotic Latent World Models

RoboJEPA:具身智慧的潛在世界模型與規模化法則

RoboJEPA is an 8B-parameter latent world model based on the JEPA architecture, establishing the first scaling laws for multi-embodiment robotic planning using real-world data.

2 min read