Aivora
arXivAI ResearchIntermediate

Building Persistent 3D Object Memory: How Ledger Tracks Objects from Egocentric Videos

打造過目不忘的 3D 空間記憶:Ledger 如何透過第一人稱影片追蹤隱形物體

2 min read
Building Persistent 3D Object Memory: How Ledger Tracks Objects from Egocentric Videos
The 30-second version

Humans recall object locations effortlessly, even without conscious intent. To give embodied assistants similar capabilities, researchers developed Ledger. It constructs a persistent 3D object memory from egocentric video streams, tracking locations, histories, and text descriptions. By clustering resting positions, Ledger filters out sensor noise and retains memory of objects even after they leave the camera's view, achieving state-of-the-art results on HD-EPIC, UCS-Bench, and Ego4D.

Key points

01

Persistent 3D Memory

Retains object locations and states in memory even after they exit the camera's field of view, including untouched items.

02

Noise-Resistant Clustering

Clusters observations by resting locations and registers movements only after repeated evidence, mitigating localization noise.

03

Rich Contextual Descriptions

Saves concise descriptions of objects (e.g., contents, supporting surfaces) to answer spatial queries without re-accessing raw videos.

04

SOTA Benchmark Performance

Improves HD-EPIC QA accuracy to 42.6% and localizes Ego4D objects with a low median error of 0.99 meters.

How it works

Ledger Persistent 3D Memory Construction Pipeline
Continuous streamFilter noiseCombine history & contextNo raw video neededEgocentric Video InputObservation TrackingResting LocationClustering3D Memory LedgerDirect Spatial QA

Why it matters

Traditional vision-language models struggle with long-term, dynamic 3D spatial memory. Ledger proves that an embodied agent can perform high-accuracy spatial retrieval using a structured '3D memory ledger' without storing massive raw video data. This paves the way for future household robots and AR smart glasses to assist users naturally, such as locating misplaced keys or remembering the contents of containers.

Who it affects

  • AI Researcher
  • AI Developer
  • Product Manager
  • Student & Learner

How to use it

  1. 1AR smart glasses assisting users in finding daily items (e.g., keys, wallets) left in other rooms.
  2. 2Home robots querying 'where is the milk?' or 'what is inside the box?' without re-processing hours of raw video.

Limitations & caveats

  • Risks of failure and performance drops in memory construction when handling long video streams stitched across multiple different scenes.
  • Relies heavily on initial 3D reconstruction and localization, which may degrade in highly dynamic or cluttered environments.

Related

IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas
arXivAI Research

IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas

IdeaAnchor:教導大語言模型將學術文獻轉化為研究點子

Researchers developed IdeaAnchor, a paradigm that trains LLMs to generate high-quality research ideas by leveraging structured specifications mined from published papers.

2 min read