Search papers, labs, and topics across Lattice.
University of California
8
0
14
LLM agents struggle with exploration in multi-agent settings, leading to poor coordination and increased regret, but a new framework can turn this around.
Current AI agents only manage to complete 20.6% of complex real-world tasks, revealing a stark gap in their capabilities compared to human users.
Local Branch Routing enables language models to leverage contextual evidence for decision-making without the computational burden of full solution searches, leading to substantial improvements in reasoning accuracy.
Misalignment in language models can be detected with 93.5% accuracy using a novel taxonomy of cognitive processes, revealing critical insights into their deceptive behaviors.
Just because your agent can write and store memories well doesn't mean it can actually *use* them effectively in a dynamic, multimodal world.
Ground-truth access in the task-generating proposer can paradoxically *accelerate* self-play collapse, suggesting that ungrounded proposers might be more stable partners for self-consistency solvers.
Forget coarse sequence-level hacks: LenVM lets you precisely dial in token generation length, boosting a 7B model's length accuracy from 30.9 to 64.8 and crushing closed-source rivals.
Even when a computer-use agent succeeds once, inconsistent task specification and variable agent behavior can tank its reliability.