QF3: Fast Flow RL with Filtered Q-Gradients
QF3:利用過濾 Q 梯度實現快速流匹配強化學習的機器人控制技術
QF3 addresses the slow training speeds of flow-based robot policies in RL. By combining flow matching with critic action gradients backpropagated through a one-step prediction, and filtering gradients to reliable action dimensions, QF3 achieves a 10x speedup over FPO++. It is the first off-policy flow RL method capable of training humanoid locomotion from scratch and transferring it zero-shot to real hardware.
Key points
Filtered Q-Gradients
Applies critic gradients only to action dimensions close to the replay buffer, maintaining update reliability.
10x Wall-Clock Speedup
Achieves a massive 10x speedup over FPO++, a recent state-of-the-art on-policy flow RL baseline.
Zero-Shot Hardware Transfer
First off-policy flow RL method to train humanoid locomotion from scratch and deploy zero-shot on physical robots.
Effective Fine-Tuning
Demonstrated strong performance in refining pre-trained flow-based manipulation policies.
How it works
Why it matters
While flow policies are highly effective for capturing complex robot actions, tuning them via RL has been notoriously slow. QF3 unlocks high-efficiency off-policy RL for flow matching. This 10x acceleration makes it highly practical to train humanoid movements from scratch and deploy them directly onto physical hardware without tuning gap issues.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Training humanoid locomotion and motion tracking from scratch
- 2Zero-shot transfer of simulation-trained policies to physical robot hardware
- 3Fine-tuning pre-trained flow-based manipulation policies for higher success rates
Limitations & caveats
- Gradient filtering heavily relies on alignment with replay actions, which may limit early exploration in highly novel state spaces.
- Performance remains bounded by the accuracy of the critic's one-step action prediction.
Related
EyeRobot 2.0: Precise Robot Manipulation via Active Gaze Without Wrist Cameras
EyeRobot 2.0:無需手腕相機,用「主動注視」實現精準雙手機器人操控
EyeRobot 2.0 mimics human vision using active gaze with a single stereo camera, enabling precise bimanual manipulation without wrist-mounted cameras.
RPG Framework: Guided Self-Improvement Boosts Embodied Agent Success to 95% Without Weight Updates
自主學習免微調!「RPG」框架藉由虛擬練習與自我診斷,將機器人任務成功率提升至 95%
The RPG framework enables autonomous robot improvement without updating model weights, boosting task success from 28.6% to 95.0% through simulation practice, failure diagnosis, and real-world transfer.
DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents
DynaHarness:具備自我演化能力的機器人代理動態實體約束框架
DynaHarness is a dynamic physical framework for self-evolving robots that bridges semantic reasoning and execution through a contract, transforming failures into capability updates.