Search papers, labs, and topics across Lattice.
5
0
8
SFT can lead to a drastic reduction in action diversity in LLMs, risking premature policy collapse even as accuracy improves.
Achieving optimal convergence rates for nonconvex optimization by transforming it into a static regret minimization problem could revolutionize how we design adaptive optimizers.
LLMs can be trained to negotiate like expert agents, extracting significantly higher surpluses by strategically exploring buyer markets rather than fixating on immediate bids.
Transformers can effectively mimic Bayesian updating processes to achieve oracle-level efficiency in average treatment effect estimation, outperforming conventional methods.
Stop rewarding all LLM-generated candidates equally: ShapE-GRPO uses Shapley values to fairly distribute credit within sets, leading to better training and faster convergence.