Long-WAM: Scaling Context for Real-Time World-Action Models
Long-WAM:突破即時控制延遲,擴展機器人世界動作模型的長上下文
This paper presents Long-WAM, a framework designed to scale visual history context for robotic world-action models under strict real-time constraints. The authors reveal that longer history only benefits control performance when the video foundation is pretrained autoregressively (AR). By first learning causal prediction from unlabeled videos and preserving this structure during action fine-tuning, combined with streaming encoding and asynchronous execution, Long-WAM achieves a low inference latency of 107.4 ms on RTX 5090. It enables highly dynamic manipulation, such as dynamic cup stacking, on physical platforms.
Key points
Autoregressive pretraining unlocks history
Simply scaling history context yields no net gain unless the video foundation is pretrained autoregressively, establishing a causal history-to-future relationship.
Longer context yields higher success
Increasing context from 0.0 to 19.2 seconds raises the success rate from 63.3% to 78.7% on RoboCasa GR-1, proving the power of temporal memory.
System optimization for real-time control
Through streaming observation encoding and asynchronous execution, Long-WAM runs at 107.4 ms per action chunk on an RTX 5090 without dropping future predictions.
Physical breakthrough in dynamic tasks
Deployed on Unitree G1 and YAM robots, achieving 95% success on dynamic cup stacking, whereas baseline methods failed completely.
How it works
Why it matters
Standard robotic policies often sacrifice visual history to maintain low latency, rendering them fragile in dynamic or long-horizon tasks. Long-WAM addresses this by demonstrating that autoregressive pretraining is essential to utilize long temporal context. By combining this insight with system-level streaming optimization, it allows robots to maintain a long memory without latency penalties, paving the way for humanoid robots to handle highly dynamic, multi-step real-world workloads.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Dynamic and rapid manipulation tasks on physical humanoid robots (e.g., dynamic cup stacking with Unitree G1).
- 2Long-horizon and multi-step composite manipulation tasks requiring consistent temporal memory.
- 3Memory-informed low-level executor complementing high-level planners in complex environments.
Limitations & caveats
- Heavy reliance on high-performance hardware to sustain sub-110ms streaming inference for action and video predictions.
- Requires large volumes of unlabeled robot and egocentric video data for the critical autoregressive pretraining phase.
Related

The Machines That Make the Machines: How NVIDIA Automates GB300 Tester Tray Assembly
機器造機器:NVIDIA 如何用 AI 與實體控制自動組裝 GB300 測試托盤
NVIDIA Seattle Robotics Lab shares insights from automating GB300 superchip tester tray assembly, highlighting the synergy between classical control engineering, smart mechanical design, and RL.
"Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models
機器人控制的「文字敏感症」:為何一個詞能讓 VLA 模型成功率從 100% 跌到 2%?
This study exposes the extreme sensitivity of Vision-Language-Action (VLA) models to instruction phrasing and introduces a zero-shot framework that uses LLMs to distill rewriting rules, significantly boosting robotic task success without retraining.
RoboJEPA: Scaling Robotic Latent World Models
RoboJEPA:具身智慧的潛在世界模型與規模化法則
RoboJEPA is an 8B-parameter latent world model based on the JEPA architecture, establishing the first scaling laws for multi-embodiment robotic planning using real-world data.