Search papers, labs, and topics across Lattice.
This paper introduces novel identifiability conditions for mixture proportion estimation (MPE) based on conditional independence (CI) given the class label, relaxing the typical irreducibility assumption. Method-of-moments estimators are derived under these CI assumptions, with asymptotic properties analyzed. The work also presents weakly-supervised kernel tests for validating the CI assumptions, demonstrating improved MPE performance and effective error control in experiments.
Identifiability in weakly supervised learning doesn't require irreducibility: conditional independence can save the day.
Mixture proportion estimation (MPE) aims to estimate class priors from unlabeled data. This task is a critical component in weakly supervised learning, such as PU learning, learning with label noise, and domain adaptation. Existing MPE methods rely on the \textit{irreducibility} assumption or its variant for identifiability. In this paper, we propose novel assumptions based on conditional independence (CI) given the class label, which ensure identifiability even when irreducibility does not hold. We develop method of moments estimators under these assumptions and analyze their asymptotic properties. Furthermore, we present weakly-supervised kernel tests to validate the CI assumptions, which are of independent interest in applications such as causal discovery and fairness evaluation. Empirically, we demonstrate the improved performance of our estimators compared with existing methods and that our tests successfully control both type I and type II errors.\label{key}