Search papers, labs, and topics across Lattice.
Affiliation:
3
0
7
0
This work introduces PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets, and introduces PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase.
Unleashing the full potential of multimodal LLMs requires reasoning directly in the visual latent space, and this paper shows how to do it with stable policy optimization.
Synthesizing realistic hand-object interactions is now possible with HO-Flow, a framework that leverages masked flow matching and interaction-aware VAEs to achieve state-of-the-art results in motion diversity and physical plausibility.