Search papers, labs, and topics across Lattice.
Peking University
2
0
4
Language models can now be rigorously evaluated on their ability to generate falsifiable research ideas, not just stylistic fluency.
By focusing on correcting "near-miss" answers, REVES achieves a remarkable +6.5 point improvement over standard RL methods, showcasing a new way to enhance LLM reasoning without extensive computational costs.