Aivora
arXivVideo AIAdvanced

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力

2 min read
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
The 30-second version

Traditional long-video AI memory frameworks rely on chronological text descriptions, which often fail to track physical identity across hours or days, confusing different objects with similar descriptions. To solve this, researchers developed Grounded Entity Biographies (GEB). GEB groups visually grounded observations of the same physical instance across clips into structured, retrievable biographies. During question answering, both the biography and episodic evidence are retrieved. GEB achieves 72.0% accuracy on EgoLifeQA, outperforming prior state-of-the-art models by 4.4 percentage points.

Key points

01

Resolving Physical Identity Confusion

Traditional timelines confuse different objects sharing similar text descriptions, whereas GEB resolves physical identity through visual grounding.

02

Constructing Entity Biographies

Groups visually grounded observations of the same physical instance across multiple long-video clips into rich, retrievable biographies.

03

Dual-Retrieval QA Mechanism

Retrieves the constructed entity biography alongside episodic evidence, enabling the model to trace objects across events during QA.

04

SOTA on Long-Video Benchmarks

Validated on four benchmarks (up to week-long videos), improving SOTA accuracy on EgoLifeQA to 72.0% (a 4.4% absolute gain).

How it works

Traditional Timeline Memory vs GEB Memory
傳統時間線記憶 (Traditional Timeline)實體自傳記憶 (GEB, Ours)
Identity Resolution依賴純文字描述,易混淆外觀相似物實體視覺定位,跨片段精準關聯同一物理對象
Memory Representation孤立、按時間線性排序的事件片段建立以物理實體為核心的專屬「自傳」
Retrieval Source僅檢索與問題關鍵字相關的事件片段同時檢索事件片段與特定實體的完整自傳
EgoLifeQA Accuracy67.6% (先前最佳)72.0% (GEB 框架)

Why it matters

This research provides key advancements for smart home assistants, egocentric wearables, and robotic monitoring. Instead of just recording chronological frames, the AI gains an understanding of how specific physical objects move and change state over days or weeks. This provides AI with true embodied long-term memory, allowing it to reason over complex, multi-day object trajectories and causal relationships.

Who it affects

  • AI Researcher
  • AI Developer
  • Product Manager

How to use it

  1. 1Egocentric personal assistants: Tracking everyday items (e.g., keys, wallet, pillbox) over days to help users find lost items or track daily routines.
  2. 2Robotic scene understanding: Enabling service robots to track and identify multiple visually similar objects over time to maintain accurate environment models.

Limitations & caveats

  • High computational overhead: Performing visual grounding and entity association across massive video scales demands high compute and memory.
  • Dependency on initial detectors: Relying heavily on front-end object detection and tracking, where early errors propagate into corrupted biographies.

Related

Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
arXivVideo AI

Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE

突破均勻分佈陷阱:透過 SplitMoE 解決影片擴散模型的擴展瓶頸

This study introduces SplitMoE, a split-role sparse architecture that bifurcates the expert pool into semantic and generic experts, overcoming the "uniformity trap" of traditional MoEs to prevent visual fragmentation in video diffusion.

2 min read