Search papers, labs, and topics across Lattice.
Nanjing University, Shanghai Innovation Institute
4
0
4
SER's innovative approach to grounding video reasoning demonstrates a 3.0-point leap in accuracy by integrating semantic verification into the reward structure.
Efficient context handling in video tasks can elevate multimodal models to new heights of agency and reasoning capability.
Teacher privilege in multimodal reasoning is redefined, showing that visually grounded cues can lead to superior performance in on-policy distillation.
Future-L1 shows that preserving visual semantics in latent space can dramatically enhance video event prediction accuracy, outperforming previous models by substantial margins.