Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong
3
0
6
Even the best LLMs fail to produce fully feasible travel plans more than half the time, revealing a significant gap in their ability to understand user needs.
Current dense predictors can deviate by over 3x from accurate predictions when faced with unsupported edge cues, revealing a fundamental flaw in their design.
Current AI agent evaluations are like testing a car only on a straight track; HAAF offers a holistic "wind tunnel" to reveal hidden risks in complex, real-world scenarios.