Search papers, labs, and topics across Lattice.
2
0
4
0
Current LLM agents struggle to keep pace with the complexities of real-life assistance, scoring low on a benchmark designed to test their proactive and persistent capabilities.
LLMs still fail at realistic search, with even the best models achieving only 30% F1 on a new benchmark designed to mimic the messy, multi-turn, and vaguely-defined search tasks that real users face.