Search papers, labs, and topics across Lattice.
4
0
5
5
Predictive divergence masks can significantly enhance RL training stability in LLMs by aligning direction criteria with actual divergence changes.
Flow-DPPO outperforms traditional PPO methods by achieving higher rewards and greater training stability through a novel divergence proximal constraint.
AffordanceVLA transforms robotic manipulation by using structured affordance cues to create precise perception-action mappings, outperforming traditional models.
Saliency-guided sparse updates, focusing on high-magnitude activations in query and key vectors, unlock significant performance gains in long-context RL, outperforming uniform update strategies.