Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
6
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.
MLLMs can revolutionize video understanding by integrating watching, remembering, and reasoning into a cohesive framework that addresses long-range dependencies and sparse evidence.
Achieving a 43.65% Effective Temporal F1 score, this work reveals that MLLMs can be effectively adapted for complex One-to-Many Temporal Grounding tasks, challenging the limitations of previous models.
Despite impressive headline scores, today's best video MLLMs can't reliably ground their answers in space and time, achieving <1% accuracy when required to identify the spatio-temporal evidence supporting their predictions.