Search papers, labs, and topics across Lattice.
5
0
6
9
Knowledge retrieval and reasoning are the main bottlenecks in KI-VQA, but a new diagnostic benchmark reveals deeper issues in visual grounding and object identification.
HiViG outperforms existing critics by integrating historical context and visual grounding, achieving up to 9% higher success rates in complex GUI tasks.
Current video generation models are far from ready to teach, stumbling on basic knowledge, skills, and attitudes needed for effective education.
Even the best search-augmented agents, like Gemini Deep Research, are easily distracted by noisy web content, leading to surprisingly poor performance (40% accuracy) on a new multimodal reasoning benchmark.
LLMs can't keep up: even state-of-the-art models struggle to adapt to dynamically changing facts in continual knowledge streams, forgetting updates and getting distracted.