Centrale LilleINRIAUniv. LilleJun 8, 2026arXiv:2606.09802

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard

AI Summary

This paper addresses the challenge of adapting multi-armed bandit algorithms to environments with user-specific preferences and drifting context distributions. The authors introduce Dri-MED, a novel algorithm that leverages linear bandit techniques to manage non-stationary heteroskedastic noise while ensuring that the mean reward exceeds a baseline strategy at each decision point. Experimental results demonstrate that Dri-MED outperforms conservative baselines by effectively accounting for user preferences and context changes, achieving a regret scaling of $\tilde{\mathcal O}\left(\fracκ{\tildeΔ}d^2(\log(T)\right)$ with manageable constraint violations.

Key Contribution

Dri-MED adapts to user preferences and context drifts, achieving significantly lower regret than traditional methods in dynamic environments.

Abstract

We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time. Under practitioner-friendly assumptions, we reduce this setting to linear bandit with stationary mean but heteroskedastic and non-stationary noise. We further study the case when the learner must ensure the mean reward of each decision must exceed that of a baseline strategy $\boldsymbolπ_0$ at each decision step. We introduce Dri-MED, an algorithm inspired from the linear version of the MED strategy, and carefully adapted to handle the non-stationary heteroskedastic noise. We show that the instance-dependent regret scales as $\tilde{\mathcal O}\left(\fracκ{\tildeΔ}d^2(\log(T)\right)$, where $\tildeΔ$ is the constraint-aware sub-optimality gap subject to policy $π_0$, with variance-aware multiplicative term $κ$ that we carefully handle using heteroskedastic regression. We further show Dri-MED enjoys $\tilde{\mathcal{O}}(d)$ expected constraint violations. Our numerical results suggest that Dri-MED significantly outperforms conservative baselines that ignores the drift and preference structure.

Recommendation & Information Retrieval RLHF & Preference Learning

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

Related Papers