Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
LLM-derived rewards can maintain optimal policy invariance even when the feedback is inaccurate, challenging the limitations of conventional reward shaping methods.