Search papers, labs, and topics across Lattice.
6
0
6
19
Fine-grained identity tuning enables precise facial edits in text-to-image models without additional training, preserving identity consistency across diverse outputs.
MV-Forcing enables the generation of long, multi-view videos with geometric consistency, overcoming the limitations of current video synthesis methods.
T2I models can be effectively probed for identity memorization without any access to training data, revealing surprising differences in how they handle famous versus less recognized names.
Executable inverse graphics can now be achieved from a single image using vision-language models, revolutionizing how we create and manipulate 3D scenes.
Floorplan localization in the wild, previously limited to controlled environments, is now robust and scalable thanks to a 3D-grounded approach that works even with single images.
Steer diffusion models to seamlessly blend pasted objects into new contexts without prompts by selectively loosening positional encoding constraints based on saliency.