Search papers, labs, and topics across Lattice.
3
1
5
5
UniPolicy is proposed, an objective-aware multi-policy alignment framework that supports parallel, business-customizable multi-policy beam search, flexibly allocating candidate quotas across objectives under a fixed retrieval budget and outperforming single-objective reinforcement learning and naive reward-fusion baselines.
Test-time RL's vulnerability to noisy pseudo-labels is amplified by group-relative advantage estimation, but can be mitigated with a surprisingly simple debiasing and denoising approach.
Forget slow and steady: "Fast Thinking" prompts, combined with carefully tuned reward functions and REINFORCE, can dramatically boost the performance of RL-trained research agents.