Search papers, labs, and topics across Lattice.
Research Center for Social Computing and Interactive Robotics Harbin Institute of Technology , China
4
0
6
2
Observation-Aligned supervision reveals that traditional chart-to-code training often leads to hallucinations, and aligning targets with identifiable quantities can dramatically improve model performance.
Even the best multimodal models struggle to reconstruct complex interactive dashboards, revealing a critical gap in current capabilities.
Language sensitivity in VLA models is a step-wise control problem, with certain task steps causing up to 50% performance degradation under non-English instructions.
Current multimodal agents are surprisingly bad at research workflows, struggling to integrate evidence across papers and figures in multi-turn settings.