Search papers, labs, and topics across Lattice.
School of Computer Science and Engineering, Northeastern University, Shenyang, China
3
0
5
2
Generative reward models can finally unlock their full potential in RL, leading to substantial performance improvements through innovative ranking strategies.
FlowCTS-OPD boosts performance metrics by over 3% while solving the temporal supervision mismatch that plagues traditional methods.
Multilingual question answering is harder than you think: even state-of-the-art RAG systems stumble when dealing with questions and knowledge in multiple languages.