For long-video question answering, visually grounding object identity across time—not just retrieving relevant clips—is critical. GEB shows that organizing observations into entity biographies improves accuracy by 4+ points on day-long and week-long videos.
This paper solves a key challenge in long-video understanding: tracking the same physical object across hours or days of footage. The authors introduce Grounded Entity Biographies (GEB), which groups visual observations of the same object into retrievable "biographies" that preserve context.