Search papers, labs, and topics across Lattice.
2
0
4
Misalignment between visual evidence and predicted timestamps in video grounding can lead to substantial performance drops, but CAVE effectively bridges this gap with boundary-specific rewards.
MLLMs can learn to reason more faithfully by explicitly anchoring visual attention to relevant image regions and reinforcing the use of that evidence during reasoning via counterfactual interventions.