Search papers, labs, and topics across Lattice.
5
0
6
8
SpanUQ reveals that span-level uncertainty quantification can enhance both the accuracy and efficiency of LLM outputs, outperforming existing methods by a factor of 10-20x.
Transforming failures into focused training tasks boosts tool-using language model performance by over 8% on key benchmarks.
ALMANAC reveals that agents can significantly improve their collaborative competence by learning from detailed human mental model annotations.
LLM agents often struggle not due to a lack of reasoning skills, but because they fail to collaborate effectively, as revealed by the new CollabSim framework.
Sustained self-improvement in LLM agents is achievable through a novel adaptive framework that outperforms traditional methods in dynamic task environments.