Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
5
Bypassing synthetic data, MLLMCLIP achieves state-of-the-art compositional accuracy in vision-language tasks through innovative feature-level distillation.
DynaVieW reveals that a schema-guided approach can drastically improve the modeling of visual dynamics, leading to superior performance in narrative generation and simulation tasks.
Current music-understanding LLMs can't tell you *when* something happens in a song, but a new benchmark and training recipe, MusTBENCH and MusT, can help them learn.