Aivora
arXivRoboticsIntermediate

"Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models

機器人控制的「文字敏感症」:為何一個詞能讓 VLA 模型成功率從 100% 跌到 2%?

2 min read
"Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models
The 30-second version

Vision-Language-Action (VLA) models suffer from extreme sensitivity to phrasing—for instance, changing "switch on the stove" to "switch on the hot plate" crashes success rates from 100% to 2%. To address this, researchers introduced "Rephrase Before You Act." By evaluating various phrasings on a few training tasks, they leverage an LLM to distill 10-20 robust rewriting rules. At runtime, incoming instructions are systematically rephrased before hitting the frozen VLA, boosting relative success rates for $π_0$ by 16% to 27% without policy retraining.

Key points

01

Extreme Language Sensitivity

A single-word change can cause success rates to swing by tens of percentage points, revealing that VLAs do not inherit the language robustness of their base VLMs.

02

Systematic Phrasing Rules

Sensitivity is highly systematic. An LLM can distill these preferences into 10 to 20 explicit rewriting rules by analyzing a small set of training tasks.

03

Zero-Shot and No Retraining

The pipeline acts as a pre-processor for instructions, requiring no retraining of the policy network and no runtime step-by-step verification.

04

Significant OOD Boost

Optimized phrasing successfully closed the 21-point performance gap between in-distribution and out-of-distribution tasks.

How it works

Rephrase Before You Act Workflow
Send raw textApply rulesOptimized phrasing (switch on the stove)Output actionsRaw Instruction (e.g.,switch on hot plate)LLM-Distilled Rules(10-20 syste…Instruction Rephraser(One-time rewrite)Frozen VLA Policy(e.g., pi_0 policy)Robotic Action (Successrate +16%-27%)

Why it matters

Traditionally, making robots robust to varied phrasing requires expensive data collection and fine-tuning. This research proves that a VLA's linguistic limitations can be bypassed by using an external LLM as an offline rule-distiller and online translator. This offers a low-cost, plug-and-play solution to make deployed robotic controllers immediately resilient to diverse, real-world human instructions.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Robotic instruction pre-processing systems that translate ambiguous human speech into highly effective instructions for VLA policies.
  2. 2Enhancing the zero-shot generalization of frozen VLA models (such as π0 and π0.5) in novel, out-of-distribution environments.

Limitations & caveats

  • The rule distillation pipeline relies on a powerful external LLM and requires initial phrasing evaluation on a subset of training tasks.
  • The synthesized rules may not fully generalize to highly idiosyncratic or extreme linguistic corner cases in unseen domains.

Related

Long-WAM: Scaling Context for Real-Time World-Action Models
arXivRobotics

Long-WAM: Scaling Context for Real-Time World-Action Models

Long-WAM:突破即時控制延遲,擴展機器人世界動作模型的長上下文

Long-WAM scales the visual history context of world-action models under real-time constraints, demonstrating that autoregressive pretraining is key to unlocking the power of long video memories.

2 min read
RoboJEPA: Scaling Robotic Latent World Models
arXivRobotics

RoboJEPA: Scaling Robotic Latent World Models

RoboJEPA:具身智慧的潛在世界模型與規模化法則

RoboJEPA is an 8B-parameter latent world model based on the JEPA architecture, establishing the first scaling laws for multi-embodiment robotic planning using real-world data.

2 min read