Search papers, labs, and topics across Lattice.
11
0
12
3
Achieving 74.8% accuracy on a new temporal reasoning benchmark, ChronoVision redefines how multimodal models can tackle complex visual tasks.
Long-horizon LLM agents can achieve 96.9% task success by learning to adapt their external execution support through trainable harness policies.
Agents can now make more accurate decisions by effectively compressing multimodal memory, closing the gap with human performance in complex environments.
DeepDebug achieves a 32% improvement in task recovery accuracy, showcasing a powerful new approach to debugging LLM agent failures.
PaperPilot transforms scientific literature search by enabling users to iteratively refine their search strategies through an interactive workflow, achieving a remarkable reduction in execution errors.
Current visual world models show a dramatic decline in performance when faced with unconventional and impossible physical interactions, highlighting a critical gap in their generalization capabilities.
Static reports are out; BioInsight's interactive system empowers researchers to dynamically explore and refine biomedical evidence like never before.
CogniRoute outperforms existing models by over 15 percentage points in social video QA, revealing the critical role of cognitive schema in multimodal reasoning.
Rationale-based fine-tuning may actually undermine clinical prediction accuracy, challenging the belief that teaching models "why" can enhance their performance.
LLMs struggle with adaptive planning, achieving only 67.75% accuracy when faced with progressively revealed world and user constraints.
LMMs can't MacGyver their way out of a paper bag: they struggle to creatively repurpose objects in visually complex environments, revealing a critical gap in grounded reasoning beyond pattern recognition.