Accelerating Video Diffusion on TPUs: Optimizing Spatio-Temporal Sparse Attention
Google TPU 影片生成加速:如何透過「時空稀疏注意力」實現 1.69 倍效能提升

High-resolution video diffusion suffers from quadratic attention scaling, where self-attention consumes up to 88.2% of per-layer latency at 1440p. Google researchers exploited the structured sparsity of Spatial-Temporal attention (SVG). By building optimized Pallas kernels on TPU v6e, they bypassed TPU hardware bottlenecks via tile-aligned boundaries and localized token permutation. This successfully achieved up to a 1.69x end-to-end speedup for 2K video generation, saving over 16 minutes per video.
Key points
Spatio-Temporal Routing
Dynamically routes attention heads into spatial or temporal masks while keeping the first frame as a global anchor.
Tile Specialization & Alignment
Separates full and boundary tiles to skip unnecessary mask evaluation, yielding a 2.40x core speedup after aligning mask boundaries.
Localized Token Permutation
Performs token permutation within device-local layouts after head exchange, avoiding expensive cross-chip communication.
High-Resolution Scaling
Provides larger savings at higher resolutions, achieving a 1.69x speedup at 1440p and saving over 16 minutes of generation time.
How it works
Why it matters
This work demonstrates how translating theoretical algorithmic sparsity into hardware speedup requires co-designing kernels and communication layouts. For developers of high-resolution video models, it proves that fine-grained hardware tailoring can save massive generation times and compute costs without sacrificing visual quality.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Accelerating high-resolution (1080p/2K) video diffusion model inference
- 2High-efficiency long-sequence self-attention computation on distributed TPU clusters
Limitations & caveats
- Aligning and rounding sparse masks slightly alters the original attention distribution, requiring a careful trade-off between video quality and speed.
- The implementation is highly specialized for TPU v6e, JAX, and Pallas Splash Attention, requiring significant rewriting for GPU architectures.
Related
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
ViTeX-Bench:解決影片場景文字編輯難題的首個高保真基準測試
Researchers introduce ViTeX-Bench and an open-source 14B model to address the challenge of editing scene text in videos while preserving temporal consistency and background locality.

NVIDIA VSS Blueprint 3.3: Lowering Visual AI Agent Costs with Smart Sampling and Agent Skills
NVIDIA VSS Blueprint 3.3 登場:以智慧採樣與 AI 代理技能,大幅降低視覺 AI 部署成本
NVIDIA VSS Blueprint 3.3 reduces development and runtime costs for visual AI agents through a prompt-based agent builder and Adaptive EVS token pruning.
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力
This paper introduces Grounded Entity Biographies (GEB), a framework that compiles cross-clip visual observations of physical objects into retrievable biographies, significantly enhancing long-video question answering.