INRIAParis-SaclayApr 15, 2026arXiv:2604.13740

Online learning with noisy side observations

AI Summary

This paper introduces a new online learning model where the learner receives noisy feedback about other actions, structured as a weighted directed graph. They define a novel graph property, the effective independence number ($α^*$), to characterize the quality of feedback. The main result is a parameter-free algorithm achieving $\widetilde{O}(\sqrt{α^* T})$ regret, improving upon existing bounds for specific partial observability models.

Key Contribution

Learning from noisy feedback doesn't have to be a guessing game: this new algorithm achieves near-optimal regret in online learning without needing to estimate the quality of the feedback.

Abstract

We propose a new partial-observability model for online learning problems where the learner, besides its own loss, also observes some noisy feedback about the other actions, depending on the underlying structure of the problem. We represent this structure by a weighted directed graph, where the edge weights are related to the quality of the feedback shared by the connected nodes. Our main contribution is an efficient algorithm that guarantees a regret of $\widetilde{O}(\sqrt{α^* T})$ after $T$ rounds, where $α^*$ is a novel graph property that we call the effective independence number. Our algorithm is completely parameter-free and does not require knowledge (or even estimation) of $α^*$. For the special case of binary edge weights, our setting reduces to the partial-observability models of Mannor and Shamir (2011) and Alon et al. (2013) and our algorithm recovers the near-optimal regret bounds.

Natural Language Processing Training Efficiency & Optimization

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Online learning with noisy side observations

Related Papers