Search papers, labs, and topics across Lattice.
This paper introduces a novel method for dimensionality reduction that enhances nearest neighbor relationships to efficiently estimate high information projections of multivariate data. By leveraging the spectral decomposition of a specially designed matrix that captures local covariance structures, the method provides a consistent estimator of the Density Information Matrix (DIM), which is crucial for tasks like Independent Components Analysis and Sufficient Dimension Reduction. The authors demonstrate that their approach not only reduces computational costs compared to existing DIM estimators but also effectively supports cluster analysis and outlier detection tasks.
Efficiently estimating high information projections could revolutionize how we approach dimensionality reduction in complex datasets.
An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data. Following similar intuitive motivation to a number of existing techniques, the proposed method is based on enhancing the nearest neighbour relationships in the data. The proposed projection arises from the spectral decomposition of a matrix designed to encode the local covariance structure in the data, where the local covariance at a point is captured by pairs of its nearest neighbours. We show that under standard regularity conditions this matrix is a consistent estimator of the so-called ``Density Information Matrix''(DIM); a non-parametric analogue of the Fisher Information Matrix. Spectral decompositions of DIMs have been shown to be connected with the important problems of Independent Components Analysis and, in the supervised context, Sufficient Dimension Reduction. However, existing estimators of the DIM are computationally expensive to compute and only target the DIM of a surrogate density, which is proportional to the square of the true underlying density. In addition, we go on to explore the practical utility of our method in aiding the downstream tasks of cluster analysis and outlier detection.