Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力
Traditional long-video AI memory frameworks rely on chronological text descriptions, which often fail to track physical identity across hours or days, confusing different objects with similar descriptions. To solve this, researchers developed Grounded Entity Biographies (GEB). GEB groups visually grounded observations of the same physical instance across clips into structured, retrievable biographies. During question answering, both the biography and episodic evidence are retrieved. GEB achieves 72.0% accuracy on EgoLifeQA, outperforming prior state-of-the-art models by 4.4 percentage points.
Key points
Resolving Physical Identity Confusion
Traditional timelines confuse different objects sharing similar text descriptions, whereas GEB resolves physical identity through visual grounding.
Constructing Entity Biographies
Groups visually grounded observations of the same physical instance across multiple long-video clips into rich, retrievable biographies.
Dual-Retrieval QA Mechanism
Retrieves the constructed entity biography alongside episodic evidence, enabling the model to trace objects across events during QA.
SOTA on Long-Video Benchmarks
Validated on four benchmarks (up to week-long videos), improving SOTA accuracy on EgoLifeQA to 72.0% (a 4.4% absolute gain).
How it works
| 傳統時間線記憶 (Traditional Timeline) | 實體自傳記憶 (GEB, Ours) | |
|---|---|---|
| Identity Resolution | 依賴純文字描述,易混淆外觀相似物 | 實體視覺定位,跨片段精準關聯同一物理對象 |
| Memory Representation | 孤立、按時間線性排序的事件片段 | 建立以物理實體為核心的專屬「自傳」 |
| Retrieval Source | 僅檢索與問題關鍵字相關的事件片段 | 同時檢索事件片段與特定實體的完整自傳 |
| EgoLifeQA Accuracy | 67.6% (先前最佳) | 72.0% (GEB 框架) |
Why it matters
This research provides key advancements for smart home assistants, egocentric wearables, and robotic monitoring. Instead of just recording chronological frames, the AI gains an understanding of how specific physical objects move and change state over days or weeks. This provides AI with true embodied long-term memory, allowing it to reason over complex, multi-day object trajectories and causal relationships.
Who it affects
- AI Researcher
- AI Developer
- Product Manager
How to use it
- 1Egocentric personal assistants: Tracking everyday items (e.g., keys, wallet, pillbox) over days to help users find lost items or track daily routines.
- 2Robotic scene understanding: Enabling service robots to track and identify multiple visually similar objects over time to maintain accurate environment models.
Limitations & caveats
- High computational overhead: Performing visual grounding and entity association across massive video scales demands high compute and memory.
- Dependency on initial detectors: Relying heavily on front-end object detection and tracking, where early errors propagate into corrupted biographies.
Related

NVIDIA VSS Blueprint 3.3: Lowering Visual AI Agent Costs with Smart Sampling and Agent Skills
NVIDIA VSS Blueprint 3.3 登場:以智慧採樣與 AI 代理技能,大幅降低視覺 AI 部署成本
NVIDIA VSS Blueprint 3.3 reduces development and runtime costs for visual AI agents through a prompt-based agent builder and Adaptive EVS token pruning.
Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
突破均勻分佈陷阱:透過 SplitMoE 解決影片擴散模型的擴展瓶頸
This study introduces SplitMoE, a split-role sparse architecture that bifurcates the expert pool into semantic and generic experts, overcoming the "uniformity trap" of traditional MoEs to prevent visual fragmentation in video diffusion.
PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching
PDMD:投影分佈匹配蒸餾技術,解決影片擴散模型的一步生成偽影
PDMD stabilizes video diffusion distillation by projecting out critic errors using the student-critic endpoint residual, achieving high-quality 4-step generation with a one-line code change.