Search papers, labs, and topics across Lattice.
This paper introduces FluxGraph, a novel approach to object transformation tracking that significantly reduces computational costs compared to the existing TubeletGraph method. By leveraging SAM2's multi-mask disagreement for selective transformation detection and eliminating the need to track all entities, FluxGraph achieves a remarkable speedup of approximately 3.3 times while enhancing tracking performance. The method demonstrates consistent efficiency improvements across multiple datasets, making it suitable for real-time applications in understanding dynamic scenes.
FluxGraph achieves over 3 times faster object tracking without sacrificing performance, revolutionizing real-time scene analysis.
Tracking objects through state transformations is essential for understanding real-world dynamics. However, existing methods are computationally expensive. TubeletGraph recently showed impressive capabilities, but its inference cost (~$4.4$ seconds per object-frame on VOST) precludes any real-time deployment possibilities. We observe that TubeletGraph's overhead arises from building a spatiotemporal partition of the input video: (1) entity segmentation is computed densely for every frame regardless of whether a transformation occurs, and (2) every entity in the scene is tracked, scaling cost with scene complexity rather than the number of transformations of interest. To address both, we propose FluxGraph, a reactive variant that uses SAM2's internal multi-mask disagreement as a lightweight trigger for transformation detection, and removes the need for tracking all entities in the given video. FluxGraph is ~$3.3\times$ faster than TubeletGraph on VOST while improving tracking performance and preserving state graph quality. Furthermore, we also observe consistent speedups of $3.7-10.7\times$ across VSCOS, M$^3$-VOS, and DAVIS17 while maintaining performance. Code is publicly available at https://github.com/YihongSun/FluxGraph.