Search papers, labs, and topics across Lattice.
Renmin University of China
2
0
3
AVOC achieves a remarkable 4.9-point accuracy boost over the next best model in long-form audio-video comprehension, redefining efficiency in multimodal understanding.
Forget disjointed pipelines and structured inputs: PlanAudio uses an LLM and semantic latent chain-of-thought to directly synthesize unified audio from free-form text prompts.