Search papers, labs, and topics across Lattice.
Affiliation:, Huawei Technologies
2
0
4
17
AI agents may ace endpoint identification but falter in delivering the evidence-based diagnostics essential for real-world telecom troubleshooting.
A-HPO significantly boosts reward acquisition in sparse-reward RL by adaptively balancing positive and negative advantage signals, outperforming GRPO, GSPO, and SAPO, especially in the critical early stages of training.