Search papers, labs, and topics across Lattice.
2
1
4
2
Over half of existing video understanding benchmarks can be solved without any visual input, exposing a critical flaw in current evaluation methods.
Object-driven shortcuts in action recognition can be effectively mitigated, leading to improved generalization in zero-shot settings.