Search papers, labs, and topics across Lattice.
2
0
3
0
MMCS achieves superior visual grounding with just 50K samples, outperforming traditional models trained on 600K image-text pairs.
Effective visual tool use hinges on causally grounded supervision, not just imitating tool calls, reshaping how we train multimodal agents.