Search papers, labs, and topics across Lattice.
6
0
8
6
AV-Flamingo outperforms existing models on complex audio-visual tasks, revealing that size isn't everything when it comes to reasoning capabilities.
Dynamic token editing in image synthesis could redefine how we approach high-resolution generative models.
A single generalist model outperforms specialized systems, achieving over 35% improvement in real-world robotic task success.
Get up to 1.79x faster ViT inference on high-resolution images without sacrificing accuracy by surgically replacing full-attention blocks with cheaper alternatives *after* pre-training.
Multimodal models can now achieve state-of-the-art performance in real-world tasks like document understanding and audio-video comprehension with significantly reduced inference latency thanks to novel token-reduction techniques.
MLLMs can now handle 4K videos up to 100x faster thanks to AutoGaze, which selectively attends to only the most informative patches.