Search papers, labs, and topics across Lattice.
School of Computer Science and Engineering, Northeastern University, Shenyang, China
3
0
5
7
Generative reward models can finally unlock their full potential in RL, leading to substantial performance improvements through innovative ranking strategies.
FlowCTS-OPD boosts performance metrics by over 3% while solving the temporal supervision mismatch that plagues traditional methods.
Reasoning with LLMs just got a whole lot faster: MemoSight cuts KV cache footprint by 66% and speeds up inference by 1.56x without sacrificing CoT performance.