Self-Correcting Multimodal Models: UMM-Reflection Enables Native Image Generation Repair via Interleaved RL
讓多模態模型自我修正!UMM-Reflection 透過交錯強化學習實現原生圖像生成反思
Unified multimodal models can theoretically repair their own image generations, but training this iterative loop is difficult. UMM-Reflection addresses this by applying RL to complete reflection trajectories inside a single unified model. It uses 'sibling trajectories' sharing an initial image to compare strategies, using a single trajectory-level advantage to update both reflection tokens and image generation. This eliminates the need for external verifiers at inference time. It improves GenEval by 12.05 points over SFT and transfers well to other unseen benchmarks like WISE and T2I-CompBench++.
Key points
Native Self-Correction Loop
A single unified model acts as both the diagnoser and reviser, iteratively observing generated images, generating reflection text, and applying pixel-level edits.
Interleaved Reinforcement Learning
Uses a single trajectory-level advantage to simultaneously optimize both text reflections and flow-based image edits, avoiding combinatorial credit assignment issues.
Sibling Trajectory Comparison
Sibling trajectories share the same starting image, enabling group-relative advantage estimation to compare and optimize different reflection strategies.
No External Verifier Needed
Once trained, the model autonomously reflects and corrects generations at inference time without relying on any external visual or textual critics.
How it works
Why it matters
Traditional image editing pipelines rely on heavy external critic-generator feedback loops. UMM-Reflection proves that a single model can natively act as both creator and critic. It improves GenEval performance by 12.05 points and shows excellent zero-shot transferability to unseen benchmarks (+10.97 on WISE), paving the way for self-improving multimodal AI agents.
Who it affects
- AI Researcher
- AI Developer
How to use it
- 1Automated image refinement and optimization
- 2Autonomous visual design and editing agents
Limitations & caveats
- Requires a supervised fine-tuning (SFT) cold start, as the model struggles to discover successful repair paths purely from scratch.
- Trajectory-level RL is computationally expensive during training due to the iterative rendering and reflection steps.
Related
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術
FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.
How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
如何循環混合專家模型?全新 Foil 架構實現專家扁平化與注意力機制解耦
This paper introduces Foil, a novel framework bridging Looped Transformers and sparse MoEs by flattening expert layers and untying attention parameters to boost training efficiency and routing quality.
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.