Search papers, labs, and topics across Lattice.
The Hong Kong Polytechnic University, Meituan Longcat Team
2
0
3
OPERA's intrinsic reward mechanism enables LLMs to achieve new heights in open-ended reasoning, outperforming proprietary models without the pitfalls of biased supervision.
LLMs can achieve massive performance gains on reasoning and knowledge-intensive tasks simply by iteratively refining their answers using pseudo-labels derived from unlabeled data.