Search papers, labs, and topics across Lattice.
The Hebrew University of Jerusalem
3
0
6
Reranking can hurt performance, but targeting high-uncertainty instances can yield significant gains while slashing computational costs.
LLM agents struggle to uncover hidden environments, with performance plummeting as task complexity increases, revealing fundamental limitations in their interactive reasoning capabilities.
Standard LLM benchmarks miss the mark: personalized "vibe-testing" reveals that user-specific prompts and subjective criteria can flip model rankings.