Search papers, labs, and topics across Lattice.
3
0
5
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Sparse feedback in auto-research could be the hidden bottleneck preventing breakthroughs, and this paper reveals how a fuzzing-inspired approach might unlock new avenues for discovery.
Progressive disclosure can significantly enhance long-context agents' performance, especially when navigating large corpora, while traditional methods falter.