Search papers, labs, and topics across Lattice.
Peking University
1
0
3
3
SpectraReward reveals that pretrained MLLMs can serve as powerful zero-shot reward models, outperforming traditional methods without the need for fine-tuning or external labels.