Search papers, labs, and topics across Lattice.
3
0
5
Key contribution not extracted.
Multimodal agents can now learn to search the web more effectively by introspecting on their own knowledge gaps, leading to a significant performance boost in visual reasoning tasks.
Multimodal agents often fail due to premature interaction collapse, but a new training scheme using structural proximity for advantage signals and differentiated Gaussian rewards can significantly improve performance.