TY - JOUR
T1 - Multi-subspace graph clustering joint dimensionality reduction and feature selection
AU - Cai, Yingjie
AU - Yang, Hui
AU - Zhu, Jianyong
AU - Nie, Feiping
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2026/4
Y1 - 2026/4
N2 - High-dimensional data present significant challenges for clustering tasks because of the curse of dimensionality and the presence of noise and excessive redundant features. Therefore, most subspace clustering methods use feature extraction techniques to project high-dimensional data into a low-dimensional space for clustering. However, they assume equal importance for all data features, ignoring their relative significance, which may not be robust enough for data noise. Additionally, dimensionality reduction, achieved by maximizing the variance of the projection space, often overlooks interpretable discriminative features in the original data. Consequently, reduced-dimensional data may lose information that is crucial for downstream clustering tasks. In this paper, we propose a novel multi-subspace graph clustering method (MGDRFS) that combines dimensionality reduction and feature selection to overcome this major problem. Unlike other single subspace clustering models, we used both dimensionality reduction and feature selection subspaces to construct the affinity matrix. Joint feature selection allows the model to focus on the most effective and interpretable discriminative features and ensures that only attributes related to the underlying manifold structure of the data are used to construct the affinity matrix, thereby effectively reducing the impact of irrelevant or noisy features and enhancing the accuracy and reliability of the affinity matrix. This synergy eliminates redundant dimensions while retaining clear and interpretable features, further improving the clustering performance and robustness of the model. To achieve this goal and overcome the challenge that the feature selection matrix is difficult to optimize under the ℓ2,0-norm constraint, we designed a new feature selection paradigm based on diagonal matrices and introduced a coordinate descent algorithm to directly solve the optimal feature selection subspace. Finally, an adaptive neighborhood strategy was adopted to construct a structured optimal similarity graph and combined it with the Laplace rank constraint to directly obtain clustering results. Numerous experimental results on real benchmark and synthetic data demonstrate that the proposed method outperforms several state-of-the-art clustering models.
AB - High-dimensional data present significant challenges for clustering tasks because of the curse of dimensionality and the presence of noise and excessive redundant features. Therefore, most subspace clustering methods use feature extraction techniques to project high-dimensional data into a low-dimensional space for clustering. However, they assume equal importance for all data features, ignoring their relative significance, which may not be robust enough for data noise. Additionally, dimensionality reduction, achieved by maximizing the variance of the projection space, often overlooks interpretable discriminative features in the original data. Consequently, reduced-dimensional data may lose information that is crucial for downstream clustering tasks. In this paper, we propose a novel multi-subspace graph clustering method (MGDRFS) that combines dimensionality reduction and feature selection to overcome this major problem. Unlike other single subspace clustering models, we used both dimensionality reduction and feature selection subspaces to construct the affinity matrix. Joint feature selection allows the model to focus on the most effective and interpretable discriminative features and ensures that only attributes related to the underlying manifold structure of the data are used to construct the affinity matrix, thereby effectively reducing the impact of irrelevant or noisy features and enhancing the accuracy and reliability of the affinity matrix. This synergy eliminates redundant dimensions while retaining clear and interpretable features, further improving the clustering performance and robustness of the model. To achieve this goal and overcome the challenge that the feature selection matrix is difficult to optimize under the ℓ2,0-norm constraint, we designed a new feature selection paradigm based on diagonal matrices and introduced a coordinate descent algorithm to directly solve the optimal feature selection subspace. Finally, an adaptive neighborhood strategy was adopted to construct a structured optimal similarity graph and combined it with the Laplace rank constraint to directly obtain clustering results. Numerous experimental results on real benchmark and synthetic data demonstrate that the proposed method outperforms several state-of-the-art clustering models.
KW - Dimensionality reduction
KW - Feature selection
KW - Multi-subspace clustering
KW - Similarity graph
UR - https://www.scopus.com/pages/publications/105020376690
U2 - 10.1016/j.patcog.2025.112557
DO - 10.1016/j.patcog.2025.112557
M3 - 文章
AN - SCOPUS:105020376690
SN - 0031-3203
VL - 172
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 112557
ER -