Search papers, labs, and topics across Lattice.
7
0
10
5
Non-thinking inference in hybrid-thinking MLLMs suffers from a staggering increase in response-pattern failures, revealing a critical misalignment that can undermine user trust.
TRACE redefines rollout budget allocation by treating each turn in a multi-turn interaction as a unique node, leading to improved reward contrast and policy performance.
Stop wasting compute on full rollouts: ADWIN dynamically adapts on-policy distillation windows, slashing training costs by up to 4.1x without sacrificing accuracy on reasoning tasks.
Most RLVR datasets are just remixes of a few originals, and this paper shows how to trace them back to their source, revealing widespread data contamination.
Training tool-calling agents with just an 8B language model outperforms traditional methods that depend on expensive resources, reshaping the landscape of tool learning.
Pass-rate-1 prompts got you down? Composition-RL boosts LLM reasoning by automatically composing multiple problems into new verifiable questions, making better use of your existing data.
Students can surpass their teachers in on-policy distillation by extrapolating rewards and merging knowledge from domain experts, challenging the conventional wisdom that students are inherently limited by their teachers' capabilities.