Search papers, labs, and topics across Lattice.
Tencent
3
0
5
RSPO transforms the training landscape for LLMs by ensuring that dense rewards enhance learning without sacrificing alignment with true outcomes.
High-entropy search tasks can now be evaluated without the constraints of incomplete ground truth, thanks to a novel framework that ensures LLMs genuinely traverse entire search spaces.
Forget bolting vision onto language models – truly powerful multimodal AI demands rethinking architectures from the ground up.