Search papers, labs, and topics across Lattice.
2
0
3
1
PCSD boosts reinforcement learning performance by 15.6 points over existing methods, demonstrating that persistent teacher signals can effectively guide agents through sparse reward landscapes.
Bridging the gap between proprietary and open-source models, MAPD achieves up to 44.4% success in QA tasks by transforming sparse RL signals into dense distillation guidance.