Search papers, labs, and topics across Lattice.
Affiliation:
5
0
6
VideoRover unifies video reasoning and external knowledge retrieval, achieving competitive performance with fewer resources than larger models.
Achieving lossless processing of 256K contexts, Keye-VL-2.0 transforms how we approach long-video understanding and agentic intelligence.
Key contribution not extracted.
Context-augmented RL lets smaller MLLMs punch *way* above their weight, rivaling much larger models on reasoning tasks while dodging reward hacking.
Achieve scalable and consistent multi-reference image editing by dynamically serializing reference images into a coherent latent sequence, outperforming existing diffusion-based methods.