Search papers, labs, and topics across Lattice.
3
0
3
6
Predictive divergence masks can significantly enhance RL training stability in LLMs by aligning direction criteria with actual divergence changes.
Smooth gradient adjustments in DRPO prevent harmful policy shifts, leading to more stable and efficient LLM training.
AgentSPEX transforms how we build and manage LLM-agent workflows, offering a modular and interpretable approach that outperforms traditional frameworks.