Search papers, labs, and topics across Lattice.
4
0
5
6
PCSD boosts reinforcement learning performance by 15.6 points over existing methods, demonstrating that persistent teacher signals can effectively guide agents through sparse reward landscapes.
CRPO effectively mitigates exposure bias in self-distillation, leading to superior performance in complex reasoning tasks.
Process evaluations reveal hidden failures in LLM reasoning, showing that lucky successes can mask critical deficiencies in agent performance.
Bridging the gap between proprietary and open-source models, MAPD achieves up to 44.4% success in QA tasks by transforming sparse RL signals into dense distillation guidance.