Search papers, labs, and topics across Lattice.
Thanks:
7
0
12
24
Group-Calibrated On-Policy Distillation boosts long-context reasoning performance by reconciling teacher guidance with verifier feedback, achieving up to a 12-point increase in benchmark scores.
Self-evolving agents are failing to adapt effectively in dynamic environments, with top methods achieving less than 70% success on benchmark tasks.
REFACT reduces token consumption while enhancing the density and faithfulness of reasoning traces in large language models, ensuring that every cited fact meaningfully supports the answer.
A simple data recipe can outperform complex reward engineering in enhancing long-context reasoning for large language models.
Larger sliding-window attention can paradoxically slow down the formation of critical retrieval mechanisms in language models, challenging conventional design assumptions.
Forget turn-based interactions: MiniCPM-o 4.5 lets you build AI that sees, hears, speaks, and *reacts* in real-time, all on a device with only 12GB of RAM.
Stop letting SFT ruin your LMMs: PRISM uses on-policy distillation to realign your model *before* RL, boosting performance by up to 6%.