Aivora
arXivImage AIAdvanced

Looped-DiT: Scaling Image Generation Efficiency via Recurrent Transformer Blocks

Looped-DiT:藉由循環 Transformer 區塊實現影像生成的高效運算擴展

2 min read
Looped-DiT: Scaling Image Generation Efficiency via Recurrent Transformer Blocks
The 30-second version

Traditional scaling for text-to-image models relies on expanding parameters or increasing denoising steps. Looped-DiT proposes a different approach: repeatedly running shared Transformer blocks within each denoising step to increase computational depth without adding parameters. To prevent the feature erosion and weak supervision of naive looping, it integrates deep supervision and self-modulating attention. This allows a 260M-parameter model to outperform a 6.5x larger model on multiple benchmarks while requiring 4.9x lower inference compute.

Key points

01

Shared Loop Mechanism

Repeatedly executes shared Transformer blocks within a single denoising step, increasing depth without adding parameters.

02

Stabilized Feature Updates

Combines deep supervision and self-modulating attention to prevent the loss of local details and address weak intermediate supervision.

03

Extreme Compute Efficiency

A 260M-parameter looped model outperforms a model 6.5x its size while requiring 4.9x lower inference compute.

04

Latent Reasoning & Correction

Increasing loop depth allows the model to progressively correct earlier generation mistakes, exhibiting latent reasoning.

How it works

Comparison of Naive Looping vs. Looped-DiT
單純循環法 (Naive Looping)Looped-DiT (本研究)
Supervision中間循環弱監督 (Weak supervision)跨中間循環深層監督 (Deep supervision)
Attention Mechanism未經調節,導致局部資訊流失 (Erodes local info)自我調節注意力,穩定特徵更新 (Self-modulating)
Image Quality無法隨深度穩定提升顯著提升,且具備潛在推理與自我糾錯能力

Why it matters

This research introduces a novel dimension for scaling generative AI. Traditionally, achieving high-quality images required deploying massive models, which creates severe memory and edge-deployment bottlenecks. Looped-DiT proves that temporal computation (looping) can substitute for spatial expansion (parameter count). This breakthrough allows high-quality visual generation to run efficiently on resource-constrained devices like smartphones without sacrificing performance.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager
  • Enterprise Leader

How to use it

  1. 1High-performance on-device text-to-image generation on resource-constrained hardware.
  2. 2Inference acceleration and deployment cost optimization for budget-constrained diffusion pipelines.

Limitations & caveats

  • While saving parameters and memory, repeatedly executing loops still requires corresponding computational time during inference.
  • Numerical stability and hyperparameter tuning at extreme loop depths might be more challenging than in non-looped models.

Related

MatLoom: Layered Text-to-Material Generation in a Compact Program Space
arXivImage AI

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

MatLoom:用極簡程式空間實現分層式文字生成 PBR 材質

MatLoom is a compact, layer-oriented language that enables LLMs to generate highly editable PBR materials as structured code, outperforming traditional diffusion models in prompt alignment.

2 min read