Search papers, labs, and topics across Lattice.
9
0
10
4
Achieving multi-stage captioning quality in a single-pass model, SimLoss FFT runs 20 times faster while maintaining high precision and recall.
Zero-shot VLMs falter at geometric reasoning, with only one out of five models surpassing random performance on basic jigsaw puzzles, revealing a scaling cliff in their capabilities.
Texture, not color, is the secret sauce behind fashion house identity, revealed by probing a multimodal CNN trained on decades of Vogue runway images.
LLMs are revolutionizing conversational AI research, and this survey offers a structured guide to navigating the rapidly evolving landscape of LLM-powered user simulation.
Multimodal agents can now plan more coherently and solve complex tasks thanks to a new anticipatory reasoning framework that forecasts short-horizon trajectories before acting.
Ditch quadratic attention in your ViTs without sacrificing performance: ViT-AdaLA distills knowledge from pre-trained VFMs into linear attention architectures, achieving state-of-the-art results on classification and segmentation.
LLMs can't keep up: even state-of-the-art models struggle to adapt to dynamically changing facts in continual knowledge streams, forgetting updates and getting distracted.
Forget direct prompt editing: this agentic planning framework, powered by offline RL and synthetic data, masters complex image styling by breaking it down into interpretable tool sequences.
Finally, AI can generate hour-long videos with consistent characters and backgrounds, thanks to a new framework that nails seamless transitions between shots.