Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
2
RL-trained multimodal models can leak sensitive information through reasoning traces, but LEMUR offers a training-free solution that effectively sanitizes this leakage without sacrificing output quality.
Transforming zero-reward training instances into valuable learning opportunities, HCGRec cuts down ineffective samples from over 70% to under 20%.
Re-ranking control alone boosts key performance metrics by over 2%, but extending this control to fine ranking yields even greater gains without compromising system stability.