Search papers, labs, and topics across Lattice.
3
0
5
5
ContextRL reveals that fine-grained context selection can lead to substantial performance boosts in LLMs, outperforming traditional data augmentation methods.
Even the top-performing conversational agents struggle with reliability, hitting only 57% accuracy on a new benchmark designed to test agentic recommender systems.
Even frontier models with high reasoning budgets fail to effectively navigate densely interlinked knowledge bases and complex policies in realistic fintech customer support scenarios, achieving only ~25.5% pass rate.