Search papers, labs, and topics across Lattice.
SenseTime Research
4
0
6
Effective visual tool use hinges on causally grounded supervision, not just imitating tool calls, reshaping how we train multimodal agents.
Authentic execution dynamics in multi-turn interactions can double the performance of OS agents, challenging the efficacy of larger models.
Ditching modular architectures unlocks surprisingly competitive vision-language performance, proving that end-to-end pixel-to-word models can rival traditional approaches at scale.
InterSketch shows that interleaving visual sketches with textual reasoning, guided by self-correction and stepwise rewards, unlocks surprisingly strong long-horizon visual reasoning, even surpassing Gemini-3-Pro.