Search papers, labs, and topics across Lattice.
3
0
4
3
Staleness-Adaptive Trust Regions reshape update geometry in asynchronous reinforcement learning, achieving record performance while controlling for high-staleness updates.
MeanFlowNFT achieves superior few-step generation performance, outperforming traditional multi-step RL-tuned methods while maintaining efficiency.
Predictive divergence masks can significantly enhance RL training stability in LLMs by aligning direction criteria with actual divergence changes.