ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
ViTeX-Bench:解決影片場景文字編輯難題的首個高保真基準測試
While video editing has progressed, modifying local 'scene text' (such as storefront signs or whiteboard text) while maintaining temporal stability remains highly challenging. This paper presents ViTeX-Bench, a benchmarking suite consisting of 387 real-world 720p videos and a 3-axis evaluation protocol spanning 13 metrics. The authors also release ViTeX-Edit-14B, a reference video editor fine-tuned with motion-aligned glyph conditioning, setting a new baseline for temporal text stability.
Key points
Real-World Dataset
Features 387 real-world 720p videos with text-region masks and editing prompts, split into 230 for training and 157 frozen for evaluation.
Three-Axis Evaluation Protocol
Evaluates text correctness, temporal quality, and edit locality via 13 metrics, supported by OCR calibration and human evaluation.
Open-Source 14B Reference Model
Releases ViTeX-Edit-14B, fine-tuned with motion-aligned glyph-video conditioning, achieving a top-tier CharAcc of 0.688.
Why it matters
Editing text embedded in video scenes (e.g., translating signs or correcting labels) is crucial for localization and VFX. Previous general metrics failed to check if text remained correct over time. ViTeX-Bench establishes a rigorous foundation for evaluating this multi-modal challenge, paving the way for reliable, localized video editing.
Who it affects
- AI Developer
- AI Researcher
- Content Creator
How to use it
- 1Video localization (translating storefront signs and posters within videos into local languages)
- 2VFX and video post-production (correcting text typos on moving product labels without reshooting)
Limitations & caveats
- Existing baseline models struggle to achieve high text accuracy, temporal stability, and background preservation simultaneously.
- The benchmark focused on 720p real videos, meaning generalization to extreme motion blur, dark settings, or higher resolutions requires further study.
Related

NVIDIA VSS Blueprint 3.3: Lowering Visual AI Agent Costs with Smart Sampling and Agent Skills
NVIDIA VSS Blueprint 3.3 登場:以智慧採樣與 AI 代理技能,大幅降低視覺 AI 部署成本
NVIDIA VSS Blueprint 3.3 reduces development and runtime costs for visual AI agents through a prompt-based agent builder and Adaptive EVS token pruning.
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力
This paper introduces Grounded Entity Biographies (GEB), a framework that compiles cross-clip visual observations of physical objects into retrievable biographies, significantly enhancing long-video question answering.
Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
突破均勻分佈陷阱:透過 SplitMoE 解決影片擴散模型的擴展瓶頸
This study introduces SplitMoE, a split-role sparse architecture that bifurcates the expert pool into semantic and generic experts, overcoming the "uniformity trap" of traditional MoEs to prevent visual fragmentation in video diffusion.