Search papers, labs, and topics across Lattice.
7
0
8
4
Recursive belief updates in AgentOPSD reveal pivotal decision points, leading to a 89.1% success rate on complex RL tasks.
Selective distillation can unlock critical learning signals in reinforcement learning, leading to significant performance gains in complex tasks.
SkillRise achieves up to 8.5 percentage points better performance than leading methods by effectively reusing transferable skills across related tasks.
Game solvers can teach LLMs how to make better decisions in long-horizon tasks by providing actionable turn-level feedback, leading to superior performance in complex environments.
MLLMs excel at single-hop tasks but falter dramatically in open-world scenarios, revealing critical gaps in their reasoning capabilities.
Asynchronous RL for LLMs doesn't have to sacrifice convergence for speed: DORA achieves 2-4x faster training by cleverly managing multiple policy versions during rollout.
Forget hand-tuning rollout budgets: $V_{0.5}$ dynamically allocates compute to sparse RL rollouts based on a real-time statistical test of a generalist value model's prior, slashing variance and boosting performance.