Search papers, labs, and topics across Lattice.
Xi鈥檃n Jiaotong University
5
0
7
CoRe achieves a remarkable 28.2-point boost in partial accuracy for cross-image reasoning, setting a new standard for vision-language models.
DataClaw_0 can transform chaotic multimodal data into structured, high-quality datasets, enhancing AI's ability to learn from less information.
Kairos achieves top-tier performance in Physical AI while ensuring efficient state management over extended time horizons, setting a new standard for operational world models.
VLMs often fail at spatial reasoning because they either ignore visual cues or exhibit unstable reasoning, but a novel process-shaping framework can fix this.
RL agents can learn more robust vision-and-language navigation policies by exploring diverse trajectories and comparing their performance, even without expert demonstrations or value networks.