Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
2
Bypassing synthetic data, MLLMCLIP achieves state-of-the-art compositional accuracy in vision-language tasks through innovative feature-level distillation.
MMHNet proves you can train a video-to-audio model on short clips and have it generalize to generate coherent audio for videos over 5 minutes long.