Search papers, labs, and topics across Lattice.
7
0
12
Scalar rewards in on-policy distillation tell tokens to gain or lose probability without specifying where that mass should actually go鈥攆raming distillation as explicit pairwise probability transport eliminates this background leakage and beats reverse-KL across reasoning tasks.
Skill contamination in LLM agents can lead to irreversible performance degradation, but a structured filtering approach can prevent this and enhance overall capabilities.
Achieving a scale error reduction to about 1% in monocular dense reconstruction could redefine the standards for visual-inertial systems.
Structured orchestration of skills, not sheer quantity, is the secret sauce for boosting LLM agent performance, with GraSP achieving remarkable efficiency gains.
Dramatically reduce hallucination in industrial RAG systems by jointly optimizing retrieval and generation with graph-aware retrieval and reinforcement learning, leading to a 92.7% reduction in URL hallucination in a real-world advertising QA system.
Agentic RAG gets a 7.7 point accuracy boost thanks to Search-P1's path-centric reward shaping, which extracts learning signals even from failed reasoning attempts.
LLMs still struggle with real-world advertising analytics, with even Gemini-3-Pro dropping to 49.4% accuracy on the most complex tasks in the new AD-Bench benchmark.