Aivora
arXivVideo AIIntermediate

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

ViTeX-Bench:解決影片場景文字編輯難題的首個高保真基準測試

2 min read
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
The 30-second version

While video editing has progressed, modifying local 'scene text' (such as storefront signs or whiteboard text) while maintaining temporal stability remains highly challenging. This paper presents ViTeX-Bench, a benchmarking suite consisting of 387 real-world 720p videos and a 3-axis evaluation protocol spanning 13 metrics. The authors also release ViTeX-Edit-14B, a reference video editor fine-tuned with motion-aligned glyph conditioning, setting a new baseline for temporal text stability.

Key points

01

Real-World Dataset

Features 387 real-world 720p videos with text-region masks and editing prompts, split into 230 for training and 157 frozen for evaluation.

02

Three-Axis Evaluation Protocol

Evaluates text correctness, temporal quality, and edit locality via 13 metrics, supported by OCR calibration and human evaluation.

03

Open-Source 14B Reference Model

Releases ViTeX-Edit-14B, fine-tuned with motion-aligned glyph-video conditioning, achieving a top-tier CharAcc of 0.688.

Why it matters

Editing text embedded in video scenes (e.g., translating signs or correcting labels) is crucial for localization and VFX. Previous general metrics failed to check if text remained correct over time. ViTeX-Bench establishes a rigorous foundation for evaluating this multi-modal challenge, paving the way for reliable, localized video editing.

Who it affects

  • AI Developer
  • AI Researcher
  • Content Creator

How to use it

  1. 1Video localization (translating storefront signs and posters within videos into local languages)
  2. 2VFX and video post-production (correcting text typos on moving product labels without reshooting)

Limitations & caveats

  • Existing baseline models struggle to achieve high text accuracy, temporal stability, and background preservation simultaneously.
  • The benchmark focused on 720p real videos, meaning generalization to extreme motion blur, dark settings, or higher resolutions requires further study.

Related

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
arXivVideo AI

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力

This paper introduces Grounded Entity Biographies (GEB), a framework that compiles cross-clip visual observations of physical objects into retrievable biographies, significantly enhancing long-video question answering.

2 min read
Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
arXivVideo AI

Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE

突破均勻分佈陷阱:透過 SplitMoE 解決影片擴散模型的擴展瓶頸

This study introduces SplitMoE, a split-role sparse architecture that bifurcates the expert pool into semantic and generic experts, overcoming the "uniformity trap" of traditional MoEs to prevent visual fragmentation in video diffusion.

2 min read