Search papers, labs, and topics across Lattice.
Affiliation:
4
1
6
0
ProgramDistill is introduced, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications, and provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.
Existing diversity metrics miss the mark, failing to capture the true variety in problem-solving strategies that could enhance LLM reasoning capabilities.
LLMs' "Aha!" moments aren't about magic tokens, but about explicitly verbalizing and managing uncertainty during reasoning, which drives performance.
LLM agents can learn to explore novel states and generalize to new tasks with a hybrid on- and off-policy RL framework that leverages memory.