Search papers, labs, and topics across Lattice.
3
0
7
1
DemoPSD effectively reduces privileged information leakage while enhancing exploration, leading to superior generalization in large language models.
Get RL-level multi-turn LLM performance with SFT-level efficiency by decoupling trajectory generation and optimization via importance weighting.
LLMs learn faster and perform better when you optimize prompts and weights together, boosting performance by 30% and cutting interaction turns by 40%.