"Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models
機器人控制的「文字敏感症」:為何一個詞能讓 VLA 模型成功率從 100% 跌到 2%?
Vision-Language-Action (VLA) models suffer from extreme sensitivity to phrasing—for instance, changing "switch on the stove" to "switch on the hot plate" crashes success rates from 100% to 2%. To address this, researchers introduced "Rephrase Before You Act." By evaluating various phrasings on a few training tasks, they leverage an LLM to distill 10-20 robust rewriting rules. At runtime, incoming instructions are systematically rephrased before hitting the frozen VLA, boosting relative success rates for $π_0$ by 16% to 27% without policy retraining.
Key points
Extreme Language Sensitivity
A single-word change can cause success rates to swing by tens of percentage points, revealing that VLAs do not inherit the language robustness of their base VLMs.
Systematic Phrasing Rules
Sensitivity is highly systematic. An LLM can distill these preferences into 10 to 20 explicit rewriting rules by analyzing a small set of training tasks.
Zero-Shot and No Retraining
The pipeline acts as a pre-processor for instructions, requiring no retraining of the policy network and no runtime step-by-step verification.
Significant OOD Boost
Optimized phrasing successfully closed the 21-point performance gap between in-distribution and out-of-distribution tasks.
How it works
Why it matters
Traditionally, making robots robust to varied phrasing requires expensive data collection and fine-tuning. This research proves that a VLA's linguistic limitations can be bypassed by using an external LLM as an offline rule-distiller and online translator. This offers a low-cost, plug-and-play solution to make deployed robotic controllers immediately resilient to diverse, real-world human instructions.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Robotic instruction pre-processing systems that translate ambiguous human speech into highly effective instructions for VLA policies.
- 2Enhancing the zero-shot generalization of frozen VLA models (such as π0 and π0.5) in novel, out-of-distribution environments.
Limitations & caveats
- The rule distillation pipeline relies on a powerful external LLM and requires initial phrasing evaluation on a subset of training tasks.
- The synthesized rules may not fully generalize to highly idiosyncratic or extreme linguistic corner cases in unseen domains.
Related

The Machines That Make the Machines: How NVIDIA Automates GB300 Tester Tray Assembly
機器造機器:NVIDIA 如何用 AI 與實體控制自動組裝 GB300 測試托盤
NVIDIA Seattle Robotics Lab shares insights from automating GB300 superchip tester tray assembly, highlighting the synergy between classical control engineering, smart mechanical design, and RL.
Long-WAM: Scaling Context for Real-Time World-Action Models
Long-WAM:突破即時控制延遲,擴展機器人世界動作模型的長上下文
Long-WAM scales the visual history context of world-action models under real-time constraints, demonstrating that autoregressive pretraining is key to unlocking the power of long video memories.
RoboJEPA: Scaling Robotic Latent World Models
RoboJEPA:具身智慧的潛在世界模型與規模化法則
RoboJEPA is an 8B-parameter latent world model based on the JEPA architecture, establishing the first scaling laws for multi-embodiment robotic planning using real-world data.