Search papers, labs, and topics across Lattice.
2
0
6
Frontier search agent performance does not require complex multi-agent swarms or test-time search verifiers: a single ReAct policy trained via iterative SFT-RL climbing hits 56.4% on Humanity's Last Exam and 92.9% on DeepSearchQA.
Because mode-seeking reverse KL aggressively amplifies incorrect teacher signals, gating dense distillation on verifier-scored teacher probes systematically outperforms uniform distillation while reclaiming massive amounts of idle teacher compute.