Search papers, labs, and topics across Lattice.
Nanjing University
5
0
9
6
Game solvers can teach LLMs how to make better decisions in long-horizon tasks by providing actionable turn-level feedback, leading to superior performance in complex environments.
RouteJudge reveals that user preferences can significantly inform the effectiveness of LLM routing strategies, transforming how we evaluate model performance.
Sparse updates in on-policy distillation can match full performance with significantly reduced training overhead, challenging conventional wisdom about dense parameter updates.
LLMs still struggle to go beyond simple lookups when answering questions about tables, especially when prediction and reasoning about unobserved data is required.
Forget hand-tuning rollout budgets: $V_{0.5}$ dynamically allocates compute to sparse RL rollouts based on a real-time statistical test of a generalist value model's prior, slashing variance and boosting performance.