Search papers, labs, and topics across Lattice.
African Institute for Mathematical Sciences
2
0
4
LLM-derived rewards can maintain optimal policy invariance even when the feedback is inaccurate, challenging the limitations of conventional reward shaping methods.
Hybrid agents that combine LLM-driven planning with RL optimization achieve superior performance in complex decision-making tasks, outperforming traditional methods.