Search papers, labs, and topics across Lattice.
University of Washington
3
0
5
9
Current vision-language models fail to achieve embodied self-awareness, with none surpassing a 16.8% success rate in real-world interaction tasks.
A unified decision process for multi-modal reasoning reveals that joint optimization of text and image generation can dramatically enhance performance in complex reasoning tasks.
LLM agents struggle to maintain performance in multi-day collaborative tasks, dropping significantly after just one environmental update, revealing a critical gap in adaptation to evolving real-world conditions.