Search papers, labs, and topics across Lattice.
4
0
5
Despite advancements in LLMs, even the best agents struggle with prospective memory, achieving only 65.1% accuracy in executing delayed intentions.
Training data diversity is the secret sauce that boosts agentic model performance, with OpenThoughts-Agent achieving a notable accuracy leap over existing benchmarks.
Generalization in LLMs hinges on training reward saturation dynamics, with reasoning faithfulness emerging as a critical predictor of success under weak supervision.
Forget brute-force scaling: targeted data curation for RLVR can unlock surprisingly large gains in LLM reasoning.