Search papers, labs, and topics across Lattice.
LIGHTSPEED
2
0
4
Skill-伪 outperforms traditional skill generation methods by leveraging a novel rollback reward mechanism, leading to significant improvements in agent performance on downstream tasks.
LLMs get stuck in their ways: even explicit corrections can't break their rigid adherence to initial (incorrect) reasoning paths in multi-turn interactions, but a novel RL approach can fix it.