Search papers, labs, and topics across Lattice.
University of California
2
0
3
Models trained on CoVA-SFT outperform existing multimodal reasoning baselines by over 2x, revealing the potential of structured visual-textual integration in AI.
Models may give clearer instructions, but they fail to engage learners deeply, resulting in passive instruction-following rather than active understanding.