Search papers, labs, and topics across Lattice.
Xi'an Jiaotong University
2
0
4
V-Zero achieves fine-grained visual reasoning without any annotated answer labels, outperforming traditional methods in both speed and accuracy.
Personality induction boosts image captioning but can hinder reasoning tasks, revealing a complex interplay in MLLM behavior that demands tailored approaches.