Search papers, labs, and topics across Lattice.
This paper introduces CrEST, a hierarchical credit assignment framework that enhances reinforcement learning with verifiable rewards (RLVR) for multi-turn tool-use agents by incorporating dense token-level supervision from a self-teacher. By addressing the challenges of inter-turn dilution and gradient concentration collapse, CrEST effectively separates credit assignment into turn-segmented verified advantages and entropy-gated self-teacher modulation. Experimental results demonstrate that CrEST significantly outperforms traditional RL and distillation methods, particularly excelling in long-trajectory tasks and strict session-level evaluations.
CrEST redefines credit assignment in RL by shifting the teacher's role from directing updates to modulating their magnitude, leading to substantial performance improvements in multi-turn agent training.
Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. $\textbf{CrEST}$ resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that $\textbf{CrEST}$ consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.