Search papers, labs, and topics across Lattice.
3
0
6
1
Conditional experience transfer can significantly enhance LLM post-training by preventing harmful updates and improving model quality.
Forget static hyperparameters: DVAO dynamically adjusts reward weights based on variance, leading to more stable and effective multi-objective RLHF.
Stop relying on stochasticity: this work shows how distilling experience into hierarchical knowledge unlocks more effective and stable LLM-based agentic search.