Search papers, labs, and topics across Lattice.
4
9
8
15
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
Task-oriented randomization can significantly enhance the stability of black-box algorithms, balancing exploration and reliability in complex input environments.
A single tokenizer, UniWeTok, now handles both high-fidelity image reconstruction and complex semantic understanding for multimodal LLMs, outperforming existing methods with far less training data.
Current LLMs and VLMs struggle with multi-step reasoning in long videos, often failing to maintain temporal coherence and procedural validity, as revealed by a new benchmark of hour-long narratives.