Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph
This paper explores the use of standard graph algorithms on UMAP's internal k-nearest-neighbor graph to enhance data analysis.
The authors demonstrate the untapped potential of UMAP's internal kNN graph for data analysis using standard graph algorithms.
Before reading this…
Applications
- →Data analysis
- →Machine learning
To understand this paper, make sure you know these concepts first:
- Understanding of UMAPfind papers →
- Familiarity with graph algorithmsfind papers →
Abstract
More Like ThisWhile UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the data manifold in its original high-dimensional space, before the distortion that UMAP's 2D projection introduces. We demonstrate the untapped potential of this internal representation, showing how standard graph algorithms applied to this graph enhance data sensemaking: (1) PageRank identifies representative data points, (2) k-core decomposition reveals dense core regions versus sparse periphery, and (3) clustering coefficient detects tight-knit neighborhoods with highly-similar data points. Through quantitative and qualitative evaluation on MNIST and Fashion MNIST, we show that these graph-based analyses are not only practical but also competitive with or complementary to purpose-built methods (e.g., k-medoids for exemplar selection, HDBSCAN for density-based clustering).