Search papers, labs, and topics across Lattice.
Affiliation:
5
0
7
3
A carefully designed critic can provide a stable and efficient alternative to traditional group-relative advantage estimation in reinforcement learning for language models.
Agents struggle to act effectively in 3D scenes, with none of the eleven evaluated VLMs achieving consistent performance across diverse tasks.
Staleness-Adaptive Trust Regions reshape update geometry in asynchronous reinforcement learning, achieving record performance while controlling for high-staleness updates.
MeanFlowNFT achieves superior few-step generation performance, outperforming traditional multi-step RL-tuned methods while maintaining efficiency.
Predictive divergence masks significantly enhance the stability of RL updates in LLMs, outperforming traditional methods by aligning direction criteria with actual divergence changes.