Search papers, labs, and topics across Lattice.
Affiliation:
10
0
12
CoRef-GS is proposed, a cooperative referring Gaussian splatting framework that constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph.
The authors', a persistent language field for structured Gaussian scenes that requires persistent semantic ownership, conserved evidence, and a hierarchy that balances stability, detail, and representation cost, is introduced.
Clip-level captions compress away the continuous visual dynamics needed for true long-horizon reasoning; Kairos restores this signal with dense, time-resolved annotations tracking actions, entities, and attributes across 10-to-30-minute videos.
X-MULTI not only synthesizes unobserved imaging factor combinations but also reveals that existing evaluation metrics can misrepresent disentanglement quality.
iARCS transforms 3D scene generation by ensuring that synthetic environments meet essential functional constraints while maintaining diversity and realism.
Get up to 20% faster ViT inference by hot-swapping certain attention heads for depthwise convolutions – without tanking accuracy.
Bridging the gap between third-person and first-person video generation is as simple as interpolating the videos, revealing that spatio-temporal discontinuities are the real bottleneck.
Autonomous vehicles can now better identify the unexpected, thanks to a new method that boosts out-of-distribution detection by up to 20% without retraining.
MLLMs can "hear" a little, but EgoSound reveals they're still largely deaf to the nuances of sound in egocentric video, especially when it comes to spatial and causal reasoning.