Search papers, labs, and topics across Lattice.
2
0
3
10
Stochastic bandit convex optimization is fundamentally harder than linear bandits, with a new lower bound that reveals a surprising increase in regret complexity.
Understanding how diverse SFT data can activate strategy selection reveals new pathways for enhancing model reasoning capabilities through targeted post-training interventions.