Search papers, labs, and topics across Lattice.
5
0
7
3
Skill selection in LLMs can be optimized to achieve a 0.73 task success rate while using 28% fewer tokens than existing methods.
Reward function design just got a major upgrade鈥擬LREF achieves 25.2% better performance by reusing and evolving reward components across iterations.
Traditional ICM strategies miss critical dynamics, leading to an average \$9,036 value error鈥擲CO redefines tournament strategy by integrating continuation values for superior outcomes.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
Unleashing an LLM's inner creativity or laser-sharp logic is now as simple as turning a knob, thanks to a new distribution-matching method that avoids heuristic rewards.