Search papers, labs, and topics across Lattice.
4
1
7
9
Over half of existing video understanding benchmarks can be solved without any visual input, exposing a critical flaw in current evaluation methods.
Reconstructing humans and their environments from multi-view video can now be done in a single pass, and 8x faster, without needing extra modules or preprocessing.
Ditch noisy feature distances: SeaCache uses a spectral filter to cache and reuse intermediate diffusion model outputs, slashing latency while maintaining image quality.
Object-driven shortcuts in action recognition can be effectively mitigated, leading to improved generalization in zero-shot settings.