Search papers, labs, and topics across Lattice.
Nanyang Technological University 鈭桬qual contribution.
3
0
6
Closing the supervision gap in GUI agents boosts success rates from the low-30% range to over 50% through innovative skill-guided learning.
Mismatched SFT data hurting your LLM's reasoning? DART uses RL to transform it into perfectly aligned training examples, boosting generalization and efficiency.
Overcome the prohibitive cost of ground-truth labels in reinforcement learning by actively acquiring labels for only the most valuable samples, leading to stable training and improved performance even with limited annotation budgets.