Search papers, labs, and topics across Lattice.
The authors tackle the open problem of meta-learning in adversarial linear contextual bandits with random action sets by formulating Meta-LinEXP3, an online-within-online framework that distills task-level parameter priors from completed tasks. This bridges a major theoretical gap by enabling provable knowledge transfer in non-stationary and worst-case reward environments without relying on stochastic assumptions. Under known context distributions, the method attains an intrinsic-dimension $\mathcal{O}(\sqrt{n})$ per-task regret bound via a policy-centered estimator, while achieving an $\mathcal{O}(n^{2/3})$ bound under unknown distributions using a past-only regularized moment estimator.
Worst-case adversarial bandits no longer have to learn from scratch: transferring predictable task-level priors across random action sets achieves provably sublinear transfer regret with intrinsic-dimension $\mathcal{O}(\sqrt{n})$ guarantees.
Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable task-level prior from completed tasks to guide the inner LinEXP3 learner. For known context distributions, we develop a policy-centered estimator that achieves an intrinsic-dimension $\mathcal{O}(\sqrt{n})$ per-task regret bound. For unknown distributions, we introduce a past-only regularized moment estimator with an $\mathcal{O}(n^{2/3})$ leading regret term and explicit finite-sample error. We further establish a direct connection between prior accuracy and transfer regret, showing that increasingly accurate priors yield sublinear transfer-dependent regret across tasks. Experiments demonstrate the effectiveness of Meta-LinEXP3, including its application to structured hyperspectral tensor sampling.