Search papers, labs, and topics across Lattice.
Stanford University
1
0
2
Systematic misalignments in MLLM-generated captions can be detected with 63.8% accuracy, revealing a critical flaw in image-text pairing that has been largely overlooked.