R4DSG: Relative 4D Scene Graph Memory for Object-Centric QA in Long Egocentric Video
A novel technique known as R4DSG (Relative 4D Scene Graph) has been developed to enhance object-centric question answering in lengthy egocentric videos. This method, outlined in a paper on arXiv (2608.11017), addresses issues such as tracking the movement of items, determining their last state change, and understanding the reasons for their relocation. While existing long-video QA approaches emphasize temporal grounding and clip retrieval, previous 3D scene-graph techniques rely on robust geometries like point clouds and RGB-D inputs, which are absent in free-motion wearable RGB video. R4DSG transforms video into compact, queryable memory entries organized by time, location, persistent objects, anchor-relative changes, and local context, thereby facilitating efficient object-centric queries. This paper suggests potential presentation at a conference or journal and is crucial for wearable AI assistants needing to respond to inquiries about object states and locations throughout extended video sequences.
Key facts
- R4DSG is a new method for object-centric question answering in long egocentric video.
- It addresses questions about where items were moved, when they changed state, and why they were relocated.
- Existing long-video QA methods emphasize temporal grounding and clip retrieval.
- Prior 3D scene-graph methods require stronger geometry such as point clouds, RGB-D, posed views, sparse reconstruction, or reconstructed scenes.
- R4DSG converts video into compact queryable memory entries indexed by time, place, persistent objects, anchor-relative change, and local interaction context.
- The method separates stable anchors from dynamic changes.
- The paper is available on arXiv with ID 2608.11017.
- The announcement type is 'cross', suggesting it may be cross-listed or presented at a conference.
Entities
Institutions
- arXiv