Search papers, labs, and topics across Lattice.
7
20
7
3
Modus achieves competitive performance across diverse benchmarks by treating all modalities symmetrically, eliminating the need for modality-specific heads or pipelines.
Current LVLMs are inadequate at fine-grained image recognition, revealing critical bottlenecks in visual and semantic processing that need urgent attention.
Track2View reduces rotation error by up to 65% and translation error by 72%, setting a new standard for 4D-consistent video generation from novel camera angles.
Achieve world-consistent video generation by directly optimizing geometry in the latent space of pre-trained video diffusion models, sidestepping costly RGB-space operations and architectural changes.
The field of video understanding is rapidly shifting from isolated pipelines to unified models capable of adapting to diverse downstream tasks, demanding a re-evaluation of current approaches.
Unlock precise, training-free color control in text-to-image models by directly manipulating the latent space's emergent Hue, Saturation, and Lightness structure.
Unlocking VLM interpretability, sparse autoencoders let you directly steer multimodal LLMs like LLaVA by intervening on CLIP's vision encoder.