Search papers, labs, and topics across Lattice.
2
0
3
0
This work proposes MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards, and shows consistent improvements across four backbones from two model families.
MERIT-Rank is proposed, a framework that models complementary reasoning trajectories to improve reranking robustness and develops Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives.