Search papers, labs, and topics across Lattice.
This paper critiques the classical data processing inequality by presenting a counterexample that demonstrates its failure in constrained learning scenarios, where the model class is limited. The authors then propose a generalized version of the inequality that establishes a lower bound on the constrained Bayes risk of a modified distribution based on a specific function set known as the superprediction set. Key findings include sufficient conditions for this containment, which could significantly influence how constrained learning problems are approached in machine learning.
Classical data processing inequality fails in constrained learning, revealing a need for a new framework to understand Bayes risk in these settings.
A key result in statistics is the data processing inequality, originally proved by Blackwell (1951) and later refined by DeGroot (1962) in terms of statistical uncertainty. It states that the Bayes risk of a statistical experiment obtained by stochastically modifying another experiment cannot be lower than the Bayes risk of the original experiment, regardless of the loss function or prior chosen. In machine learning, this result underlies applications such as the information bottleneck principle and some feature learning techniques. However, machine learning problems are constrained learning problems: the model class used does not include all measurable functions. We present a simple counterexample showing that the classical data processing inequality fails to hold in such a setting. Hence, we formulate a generalized data processing inequality, requiring the constrained Bayes risk of a joint distribution (with respect to a loss function and a constrained hypothesis class) to lower bound the constrained Bayes risk on the stochastically modified distribution, regardless of the choice of distribution. We show this inequality to be equivalent to a set containment condition on a specific function set induced by the loss and model class, called the superprediction set. Finally, we derive sufficient conditions for this containment.