Aivora
Google AI DevelopersVideo AIAdvanced

Accelerating Video Diffusion on TPUs: Optimizing Spatio-Temporal Sparse Attention

Google TPU 影片生成加速:如何透過「時空稀疏注意力」實現 1.69 倍效能提升

2 min read
Accelerating Video Diffusion on TPUs: Optimizing Spatio-Temporal Sparse Attention
The 30-second version

High-resolution video diffusion suffers from quadratic attention scaling, where self-attention consumes up to 88.2% of per-layer latency at 1440p. Google researchers exploited the structured sparsity of Spatial-Temporal attention (SVG). By building optimized Pallas kernels on TPU v6e, they bypassed TPU hardware bottlenecks via tile-aligned boundaries and localized token permutation. This successfully achieved up to a 1.69x end-to-end speedup for 2K video generation, saving over 16 minutes per video.

Key points

01

Spatio-Temporal Routing

Dynamically routes attention heads into spatial or temporal masks while keeping the first frame as a global anchor.

02

Tile Specialization & Alignment

Separates full and boundary tiles to skip unnecessary mask evaluation, yielding a 2.40x core speedup after aligning mask boundaries.

03

Localized Token Permutation

Performs token permutation within device-local layouts after head exchange, avoiding expensive cross-chip communication.

04

High-Resolution Scaling

Provides larger savings at higher resolutions, achieving a 1.69x speedup at 1440p and saving over 16 minutes of generation time.

How it works

Token Permutation Pipeline for Reducing Communication Overhead in Distributed Inference
Cross-chipHead-localTemporal layoutRestore F,H,WOutput exchangeDistributed Heads InputAll-to-All ExchangeLocal PermutationSparse AttentionRestore Token OrderReturn Exchange

Why it matters

This work demonstrates how translating theoretical algorithmic sparsity into hardware speedup requires co-designing kernels and communication layouts. For developers of high-resolution video models, it proves that fine-grained hardware tailoring can save massive generation times and compute costs without sacrificing visual quality.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Accelerating high-resolution (1080p/2K) video diffusion model inference
  2. 2High-efficiency long-sequence self-attention computation on distributed TPU clusters

Limitations & caveats

  • Aligning and rounding sparse masks slightly alters the original attention distribution, requiring a careful trade-off between video quality and speed.
  • The implementation is highly specialized for TPU v6e, JAX, and Pallas Splash Attention, requiring significant rewriting for GPU architectures.

Related

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
arXivVideo AI

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

ViTeX-Bench:解決影片場景文字編輯難題的首個高保真基準測試

Researchers introduce ViTeX-Bench and an open-source 14B model to address the challenge of editing scene text in videos while preserving temporal consistency and background locality.

2 min read
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
arXivVideo AI

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力

This paper introduces Grounded Entity Biographies (GEB), a framework that compiles cross-clip visual observations of physical objects into retrievable biographies, significantly enhancing long-video question answering.

2 min read