Search papers, labs, and topics across Lattice.
Griffith University
2
0
3
Categorical value learning can significantly enhance the performance of PPO critics in reinforcement learning, leading to better calibration and lower variance in advantage estimation.
Qwen-UI-Agent outperforms leading models in mobile and cross-platform tasks, achieving up to 97.5% accuracy on key benchmarks.