摘要
Although flat and hierarchical clustering methods have been widely used, both are limited in their ability to model multi-resolution data structures. Flat clustering operates at a single resolution and requires the number of clusters to be predefined, whereas hierarchical clustering provides a full hierarchy but typically suffers from high computational cost and lacks principled mechanisms for selecting meaningful resolutions. In this work, we propose an Efficient Hierarchical k-means (EHKM) algorithm to enable scalable clustering across multiple resolutions. EHKM starts by treating each data point as an individual cluster and each cluster is merged with the one that minimizes the increase in the sum of squared errors (SSE) at the same time, thereby aligning with a greedy optimization of the k-means objective. This process can be repeated, gradually adjusting the clustering resolution from fine to coarse. To efficiently obtain clusterings with an arbitrary number, we introduce a refinement procedure based on a lazy-update scheme. Our theoretical analysis and experimental results show that the number of clusters in the hierarchy decreases exponentially, ensuring high computational efficiency. Experiments on real-world datasets demonstrate that EHKM achieves competitive clustering performance with significantly improved scalability, making it a practical tool for uncovering meaningful structures in large-scale multi-granularity data.
| 源语言 | 英语 |
|---|---|
| 期刊论文编号 | 112930 |
| 期刊 | Pattern Recognition |
| 卷 | 174 |
| DOI | |
| 出版状态 | 已出版 - 6月 2026 |
学术指纹
探究 'Efficient hierarchical multi-resolution k-means clustering' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver