Search papers, labs, and topics across Lattice.
This paper introduces D'茅j脿 Cue, a training-free framework that enhances the localization of object states in visual histories by utilizing a vocabulary-relative coordinate system. By leveraging alternative state descriptions, the method effectively identifies intervals where specific states hold, significantly improving retrieval accuracy on object histories. The results demonstrate a nearly twofold increase in retrieval performance, with R@1 at tIoU 0.5 rising from 10.3% to 20.5% on 78 VOST histories, highlighting the effectiveness of using state-balanced centroids for calibration.
Vocabulary-relative queries can nearly double the accuracy of state localization in object histories, transforming how we understand visual tracking.
Tracking links observations of the same object through visual change, yet cannot by itself determine when the object is empty or filled, intact or cut. We formulate identity-conditioned state-moment retrieval: given a tracked-object history and alternative state descriptions, localize an interval in which each described state holds. Absolute image-text similarity scores descriptions independently; because every visible frame depicts the same target, shared object compatibility can obscure the state evidence needed to identify the target interval. The alternatives provide the missing reference: evidence for one state should be measured against the others. We introduce D\'ej\`a Cue, a training-free framework that turns these alternatives into a vocabulary-relative coordinate system. It subtracts their state-balanced centroid from each description, calibrates frame scores, and scans multiple durations within contiguous visible runs using a frozen encoder. On 78 VOST histories, holding the temporal scan fixed and changing only the query reference nearly doubles R@1 at tIoU 0.5 from 10.3\% to 20.5\% and raises Top-1 tIoU from 16.0\% to 21.5\%. Candidate-rank analyses show that vocabulary-relative queries rank useful intervals higher within the same candidate set. Related state descriptions can therefore serve as an object-specific, query-time coordinate system for reading frozen visual representations.