Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
2
0
1
Regret and instability in multi-armed bandits are fundamentally intertwined, with a new algorithm that optimally balances both while matching established lower bounds.
A non-adaptive protocol can achieve order-optimal one-bit mean estimation without any interaction, challenging long-held assumptions in the field.