Search papers, labs, and topics across Lattice.
Politecnico di Milano
3
0
5
Transforming long-form videos into compact, temporally grounded scene graphs allows MLLMs to maintain semantic richness while overcoming input constraints, leading to state-of-the-art VQA performance.
Current multimodal anomaly detection models are misrepresented by existing benchmarks, showing only superficial reliance on text guidance in decision-making.
Finally, a feed-forward method cracks dynamic 3D scene reconstruction from multi-view video without needing camera poses, opening the door to real-time applications.