Search papers, labs, and topics across Lattice.
Nvidia
3
0
6
18
The accuracy gap in multilingual reasoning can swing dramatically by up to 57 points based on output-token caps, challenging conventional evaluation methods.
Human-written policies can boost agent performance significantly, but learning from experience remains a major hurdle for effective text policy generation.
Structured feedback can boost LLM agent success rates by up to 44 percentage points, revealing the critical role of admissible alternatives in the repair process.