Search papers, labs, and topics across Lattice.
2
0
3
1
OmniScope reveals that treating audio and video relevance separately can drastically enhance performance in omnimodal models, achieving remarkable efficiency gains without sacrificing accuracy.
Leaderboard-topping video models are still surprisingly brittle, failing on basic video reasoning tasks unless given the right textual cues.