Search papers, labs, and topics across Lattice.
Affiliation:
4
0
8
Shared reasoning structures can significantly boost performance in continual reinforcement learning, with a novel replay mechanism achieving parity with multitask training.
Even advanced LLMs struggle with budget-constrained combo shopping, revealing a significant gap in their ability to handle real-world constraints effectively.
MoEs, despite their scaling advantages, suffer from a surprising "spectral plasticity loss" in continual RL, but a simple Parseval penalty can recover performance.
Forget hand-crafted reward functions: MVR uses multi-view video and a frozen VLM to automatically shape RL rewards, teaching agents complex motions without getting stuck on static poses.