Search papers, labs, and topics across Lattice.
6
0
10
34
Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both effects.
Short-context models can achieve superior reasoning performance by leveraging long-context teacher models through innovative token alignment and training strategies.
Draft-OPD accelerates inference by over 5x while improving speculative decoding accuracy, transforming how draft models learn from target feedback.
Test-time training can finally scale for large reasoning models: TEMPO unlocks sustained performance gains by interleaving policy refinement with periodic critic recalibration, boosting accuracy by over 18% on challenging benchmarks.
A new 32B code LLM trained specifically for industrial tasks crushes existing models on specialized domains like chip design and GPU kernel optimization, while remaining competitive on general coding benchmarks.
Intrinsic reward signals in unsupervised RL for LLMs inevitably collapse due to sharpening of the model's prior, but external rewards grounded in computational asymmetries offer a path to sustained scaling.