Search papers, labs, and topics across Lattice.
JD.com Inc
1
0
2
Naive linear reward scalarization inevitably triggers destructive gradient tug-of-wars, but decoupling marginal objectives via orthogonal projections and localized constraints stabilizes training across both recommender systems and complex LLM alignment tasks.