Search papers, labs, and topics across Lattice.
2
0
3
AVSCap-7B achieves superior audio-visual synergy, outperforming existing models by effectively linking non-speech sounds to visual actions.
Open-weight Omni models struggle with binding accuracy, achieving only 41.55% on a new counterfactual benchmark, highlighting a critical gap in long-video comprehension.