Search papers, labs, and topics across Lattice.
Sony Group Corporation
2
0
4
1
Current music-understanding LLMs can't tell you *when* something happens in a song, but a new benchmark and training recipe, MusTBENCH and MusT, can help them learn.
MMHNet proves you can train a video-to-audio model on short clips and have it generalize to generate coherent audio for videos over 5 minutes long.