Search papers, labs, and topics across Lattice.
2
0
4
Current vision-language models falter in streaming interaction understanding, with alarming mis-calibration leading to confidently incorrect predictions.
MLLMs struggle to juggle proactive tasks and reactive queries in dynamic video streams, but a simple agentic framework can significantly improve their coordination without any training.