Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
4
Gaussian guidance can enhance reinforcement learning efficiency by adapting trajectory retention depth dynamically, leading to substantial performance gains at reduced costs.
Relay-OPD not only corrects reasoning missteps in real-time but also slashes training time by over half, setting a new standard for on-policy distillation efficiency.
LLMs can learn to recognize when they lack sufficient information for reasoning and proactively ask for clarification, leading to more reliable and concise answers.