Search papers, labs, and topics across Lattice.
This paper critically examines existing evaluation techniques for unsupervised feature selection, revealing that many are fundamentally flawed and rely on supervised elements. The authors propose a novel evaluation framework that leverages unsupervised Principal Component Analysis and optimal transport to assess feature selection quality without any label information. Their approach establishes a truly unsupervised methodology, enhancing the integrity of feature selection evaluations in data mining.
Current unsupervised feature selection evaluations are often misleading, masking supervised influences that compromise their validity.
Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method. Most of the methods commonly used for the unsupervised evaluation of feature selection algorithms suffer from critical design flaws which question their unsupervised nature. In this paper, we provide a critical discussion on the established allegedly unsupervised evaluation techniques, and shed light on the reasons why they are not truly unsupervised but, at best, supervised evaluation under an unsupervised downstream task. We also propose a novel, truly unsupervised evaluation framework to measure the quality of the feature selection algorithms without any form of information about the labels. The proposed framework utilizes unsupervised Principal Component Analysis, and optimal transport to measure the quality of the feature selection methods in a truly unsupervised manner.