Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
0
Behavior policies trained in simulation can be robust to real-world transition uncertainties, reducing evaluation variance and reliance on costly real-world samples.
LLMs can be trained to negotiate like expert agents, extracting significantly higher surpluses by strategically exploring buyer markets rather than fixating on immediate bids.
Ditching pessimism unlocks a quadratically faster, $\widetilde{\mathcal{O}}(1/n)$ sample complexity for offline learning in KL-regularized games.