Looped-DiT: Scaling Image Generation Efficiency via Recurrent Transformer Blocks
Looped-DiT:藉由循環 Transformer 區塊實現影像生成的高效運算擴展
Traditional scaling for text-to-image models relies on expanding parameters or increasing denoising steps. Looped-DiT proposes a different approach: repeatedly running shared Transformer blocks within each denoising step to increase computational depth without adding parameters. To prevent the feature erosion and weak supervision of naive looping, it integrates deep supervision and self-modulating attention. This allows a 260M-parameter model to outperform a 6.5x larger model on multiple benchmarks while requiring 4.9x lower inference compute.
Key points
Shared Loop Mechanism
Repeatedly executes shared Transformer blocks within a single denoising step, increasing depth without adding parameters.
Stabilized Feature Updates
Combines deep supervision and self-modulating attention to prevent the loss of local details and address weak intermediate supervision.
Extreme Compute Efficiency
A 260M-parameter looped model outperforms a model 6.5x its size while requiring 4.9x lower inference compute.
Latent Reasoning & Correction
Increasing loop depth allows the model to progressively correct earlier generation mistakes, exhibiting latent reasoning.
How it works
| 單純循環法 (Naive Looping) | Looped-DiT (本研究) | |
|---|---|---|
| Supervision | 中間循環弱監督 (Weak supervision) | 跨中間循環深層監督 (Deep supervision) |
| Attention Mechanism | 未經調節,導致局部資訊流失 (Erodes local info) | 自我調節注意力,穩定特徵更新 (Self-modulating) |
| Image Quality | 無法隨深度穩定提升 | 顯著提升,且具備潛在推理與自我糾錯能力 |
Why it matters
This research introduces a novel dimension for scaling generative AI. Traditionally, achieving high-quality images required deploying massive models, which creates severe memory and edge-deployment bottlenecks. Looped-DiT proves that temporal computation (looping) can substitute for spatial expansion (parameter count). This breakthrough allows high-quality visual generation to run efficiently on resource-constrained devices like smartphones without sacrificing performance.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
- Enterprise Leader
How to use it
- 1High-performance on-device text-to-image generation on resource-constrained hardware.
- 2Inference acceleration and deployment cost optimization for budget-constrained diffusion pipelines.
Limitations & caveats
- While saving parameters and memory, repeatedly executing loops still requires corresponding computational time during inference.
- Numerical stability and hyperparameter tuning at extreme loop depths might be more challenging than in non-looped models.
Related
Embedding Prediction Helps Image Generation: NEPA-DiT Achieves Superior FID with Much Less Compute
預測嵌入向量助力影像生成:全新 NEPA-DiT 框架以超低算力達到更優 FID
This study introduces NEPA, a framework that dynamically predicts and updates embedding conditions during denoising, enabling NEPA-DiT-XL to achieve an impressive 1.32 FID with only one-third of REPA's training compute.
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
MatLoom:用極簡程式空間實現分層式文字生成 PBR 材質
MatLoom is a compact, layer-oriented language that enables LLMs to generate highly editable PBR materials as structured code, outperforming traditional diffusion models in prompt alignment.
MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation
統一分佈訓練框架 MGFlow:實現高品質單步視覺生成
This paper presents a unified theoretical framework for distributional training and introduces MGFlow, which uses Gaussian mixtures and Wasserstein gradient flows to achieve state-of-the-art one-step visual generation.