Search papers, labs, and topics across Lattice.
RISE introduces a comprehensive framework for understanding roadside infrastructure through a combination of metric 3D tracking and structured vision-language reasoning. The method utilizes an image-only approach that leverages SAM3 video identities and calibration-guided mask agreement to achieve persistent 3D tracks across multiple camera views without the need for LiDAR or specific 3D training, achieving a MOTA of 66.9 on a dataset of human-reviewed clips. Additionally, the framework includes the RISE-VQA dataset, which facilitates advanced QA tasks by providing 33,910 pairs derived from diverse roadside scenarios, highlighting the importance of domain adaptation and temporal context in improving performance while addressing challenges in spatial grounding and interaction reasoning.
Achieving 66.9 MOTA in 3D tracking without LiDAR reveals a new frontier in roadside infrastructure understanding.
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.