Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
Visual blind spots can be transformed into powerful self-supervision signals, leading to substantial performance gains in multimodal language models.
Visual under-conditioning in LMMs can be overcome by directly regularizing visual attention, leading to remarkable improvements in multimodal understanding tasks.
Unsupervised training can elevate multimodal models, achieving a 3.5% boost in understanding metrics and a notable increase in image generation fidelity without human intervention.