Search papers, labs, and topics across Lattice.
Affiliation:
5
0
7
12
Disentanglement and attribute binding are the real bottlenecks in multi-reference image generation, not scene composition, with top models still struggling to achieve high fidelity.
Current video generation models fail to maintain embodiment consistency and functional interaction in human-to-robot manipulation, revealing significant gaps in their transfer capabilities.
Current vision-language models struggle with process understanding in robotic manipulation, but targeted post-training can yield significant improvements.
Revisiting spatial reasoning allows models to correct initial hypotheses with new perspectives, dramatically enhancing their accuracy in complex environments.
Ditch the feature extraction pipeline: GenMask directly generates segmentation masks with a diffusion transformer, achieving SOTA results by harmonizing mask and image generation in a single model.