Search papers, labs, and topics across Lattice.
2
0
4
MoVA achieves superior video-text alignment by disentangling evolving visual concepts from static textual descriptions, outperforming existing models in handling long sequences.
Generalist agents need to remember distinct information to navigate conflicting optimal actions across environments, challenging the notion that current state observations are sufficient.