Search papers, labs, and topics across Lattice.
3
0
5
Achieving up to 48.3% better image quality and 8.9% higher accuracy in face recognition, MTVDiff revolutionizes thermal-to-visible face translation through advanced multimodal integration.
Generating realistic human-object interaction videos from text, images, audio, *and* pose is now possible, opening the door to automated content creation workflows.
Analytical diffusion models can now scale to ImageNet-1K without training, thanks to a clever "Golden Subset" selection strategy that avoids full-dataset scans.