Search papers, labs, and topics across Lattice.
2
0
3
RepFusion reveals that multimodal large language models can dramatically enhance denoising in text-to-image systems, outperforming traditional denoising methods.
Camera pose, largely ignored in video LLMs, unlocks significant gains in spatial reasoning and even improves general video QA when used as a lightweight supervisory signal.