Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Gaussian guidance can enhance reinforcement learning efficiency by adapting trajectory retention depth dynamically, leading to substantial performance gains at reduced costs.
LLMs can escape the trap of confidently wrong reasoning by co-evolving a generator and verifier from a single model, bootstrapping each other to break free from flawed consensus.