Search papers, labs, and topics across Lattice.
Princeton University
5
1
9
11
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
Targeted middle-layer recurrence can dramatically enhance Transformer reasoning without the need for full-layer looping, leading to superior performance in complex tasks.
Privileged self-distillation can paradoxically hinder thinking models, leading to a 17% drop in accuracy on long reasoning tasks due to its impact on learning dynamics.
Agentic coding gets a serious boost: distilling and reusing rollout trajectories lets Claude-4.5-Opus jump from 70.9% to 77.6% on SWE-Bench Verified.
Turn sparse binary rewards into dense supervision signals by having a model revise its own work, then distilling the revision strategy back into the original generation.