Search papers, labs, and topics across Lattice.
1
0
2
DSCC is the first method to effectively ground long-form captions in multimodal models by integrating visual anchors during training, achieving unprecedented precision and length in caption generation.