Search papers, labs, and topics across Lattice.
Affiliation:
2
0
2
VIBE achieves superior instruction adherence and controllability in music generation from video, setting a new standard for semantic alignment in multimodal tasks.
TEMPO not only achieves superior timestamping accuracy in audio-language models but also integrates innovative techniques like atomic timestamp tokens and reinforcement learning for enhanced performance.