Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong
2
0
4
Even the best LLMs fail to produce fully feasible travel plans more than half the time, revealing a significant gap in their ability to understand user needs.
Current AI agent evaluations are like testing a car only on a straight track; HAAF offers a holistic "wind tunnel" to reveal hidden risks in complex, real-world scenarios.