Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
Tracking both task advantage and actual reward gains can drastically improve the efficiency of multi-task reinforcement learning for LLMs.
T3S achieves unprecedented efficiency in multi-task reinforcement learning by tailoring features and task selection, outperforming existing methods in robotics tasks.
LLM memory failures are systematic, stemming from operation-level issues like information loss and retrieval misalignment, and can be automatically corrected with prompt optimization guided by fine-grained error tracing.