Search papers, labs, and topics across Lattice.
This paper explores semi-supervised classification using a two-component Weibull mixture model, focusing on the impact of informative missing labels on classification performance. By modeling the probability of missing labels as a function of classification uncertainty, the authors establish a feature-dependent missing-at-random mechanism that enhances the classifier's performance. Key findings include the characterization of decision regions and the derivation of asymptotic relative efficiency formulas, demonstrating significant potential for reducing expected error rates through improved decision-boundary estimation.
Informative missing labels can significantly enhance classification accuracy by providing insights into uncertainty, leading to reduced expected error rates in semi-supervised settings.
We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent missing-at-random (MAR) mechanism that shares parameters with the Weibull-mixture classifier. The missing-label indicators can therefore provide information about the classifier in addition to the observed features and available class labels. Under a common Weibull shape, a Bayes' rule has at most one positive decision boundary, which is unique when the rule is nonconstant; under unequal shapes, it can have two. We characterise these decision regions, derive the Fisher information for the classifier after adjustment for nuisance parameters in the missingness model, and obtain a decision-boundary expansion of the expected error rate of the plug-in sample rule relative to the Bayes error. The expansion yields classification-specific asymptotic relative efficiency formulas for the one- and two-boundary cases and shows that a positive-definite increase in Fisher information is sufficient, but not necessary, for a smaller first-order expected error rate. Numerical studies and a semi-synthetic analysis based on hard-drive failure data illustrate potential reductions in expected error rate and improvements in decision-boundary estimation from modelling feature-dependent label missingness.