Search papers, labs, and topics across Lattice.
5
0
6
6
MLLMs can encode visual evidence but often fail to control their reliance on it, revealing a critical bottleneck in multimodal reasoning.
Modus achieves competitive performance across diverse benchmarks by treating all modalities symmetrically, eliminating the need for modality-specific heads or pipelines.
Track2View reduces rotation error by up to 65% and translation error by 72%, setting a new standard for 4D-consistent video generation from novel camera angles.
Achieve world-consistent video generation by directly optimizing geometry in the latent space of pre-trained video diffusion models, sidestepping costly RGB-space operations and architectural changes.
Imagine designing custom fonts simply by describing them or providing a reference image – VecGlypher makes it a reality by directly generating editable vector glyphs with a single multimodal language model.