Search papers, labs, and topics across Lattice.
Carnegie Mellon University
4
8
9
10
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.
Forget expensive per-task search: agentic workflows can be synthesized in a single LLM pass by transferring learned structural priors, slashing optimization costs by 3 orders of magnitude.
Forget manual curation鈥攁ligning policy gradients with a validation set adaptively selects RL training data, leading to more stable LLM training and improved performance.
Forget context window limits: this RL method uses LLM-generated summaries to train agents for long-horizon tasks, achieving higher success rates with less context.