TY - JOUR
T1 - Learning Compact Discriminant Representation via Low-Rank Bilinear Pooling
AU - Song, Kun
AU - Li, Hao
AU - Cheng, Gong
AU - Han, Junwei
AU - Nie, Feiping
AU - Gu, Bin
AU - Karray, Fakhri
N1 - Publisher Copyright:
© 1979-2012 IEEE.
PY - 2025
Y1 - 2025
N2 - In this paper, we explain the mechanism of bilinear pooling as a module of hard sample generation, and find that bilinear pooling significantly expands variances of the first-order vectors when it produces discriminative bilinear features. In conjunction with the extremely high dimensionality of the obtained bilinear features, those variances lead to overfitting in subsequent learning models. To solve this issue, we construct a bi-level optimization problem, where the high-level problem is the supervised classification loss, and the low-level problem is the principal component analysis (PCA). Then, we find that PCA on bilinear features is equivalent to spectral clustering, which allows us to mathematically prove that the first log 2(C) principal components can support the discriminant information of C classes. By removing the rest principal components, the dimensionality and variances are simultaneously reduced. To the best of our knowledge, this is the first work providing a lower bound for dimension reduction for bilinear pooling. However, the PCA projection matrix L is prone to overfitting due to having many parameters. To address this issue, we propose a rank-k general bilinear projection (RK-GBP) that decomposes L into two small matrices U and V, whose learnable parameters are smaller. Different from traditional bilinear projections used in factorized bilinear pooling (FBiP), our RK-GBP can preserve the orthogonality of columns in L by constraining the orthogonality of columns in Uand V. For computational efficiency, we relax the PCA in the low-level task into a dictionary learning problem, obtaining the rank-k orthogonal factorization bilinear pooling (RK-OFBP). The RK-OFBP can be considered as a general form of current factorization bilinear pooling methods (e.g., Hadamard product-based ones). Finally, we evaluate our approach on fine-grained images and large-scale datasets, demonstrating that our proposed method not only produces extremely low-dimensional features but also outperforms other methods in classification tasks. For example, our RK-OFBP can employ 32-dimensional vectors to achieve comparable results to B-CNN (Lin, 2015) (dimension: 512*512) for the 200-class classification task.
AB - In this paper, we explain the mechanism of bilinear pooling as a module of hard sample generation, and find that bilinear pooling significantly expands variances of the first-order vectors when it produces discriminative bilinear features. In conjunction with the extremely high dimensionality of the obtained bilinear features, those variances lead to overfitting in subsequent learning models. To solve this issue, we construct a bi-level optimization problem, where the high-level problem is the supervised classification loss, and the low-level problem is the principal component analysis (PCA). Then, we find that PCA on bilinear features is equivalent to spectral clustering, which allows us to mathematically prove that the first log 2(C) principal components can support the discriminant information of C classes. By removing the rest principal components, the dimensionality and variances are simultaneously reduced. To the best of our knowledge, this is the first work providing a lower bound for dimension reduction for bilinear pooling. However, the PCA projection matrix L is prone to overfitting due to having many parameters. To address this issue, we propose a rank-k general bilinear projection (RK-GBP) that decomposes L into two small matrices U and V, whose learnable parameters are smaller. Different from traditional bilinear projections used in factorized bilinear pooling (FBiP), our RK-GBP can preserve the orthogonality of columns in L by constraining the orthogonality of columns in Uand V. For computational efficiency, we relax the PCA in the low-level task into a dictionary learning problem, obtaining the rank-k orthogonal factorization bilinear pooling (RK-OFBP). The RK-OFBP can be considered as a general form of current factorization bilinear pooling methods (e.g., Hadamard product-based ones). Finally, we evaluate our approach on fine-grained images and large-scale datasets, demonstrating that our proposed method not only produces extremely low-dimensional features but also outperforms other methods in classification tasks. For example, our RK-OFBP can employ 32-dimensional vectors to achieve comparable results to B-CNN (Lin, 2015) (dimension: 512*512) for the 200-class classification task.
KW - Rank-k bilinear projection
KW - bilinear pooling
KW - linear dimensionality reduction
KW - normalized cuts
KW - spectral clustering
UR - https://www.scopus.com/pages/publications/105013997248
U2 - 10.1109/TPAMI.2025.3601355
DO - 10.1109/TPAMI.2025.3601355
M3 - 文章
AN - SCOPUS:105013997248
SN - 0162-8828
VL - 47
SP - 10914
EP - 10931
JO - IEEE Transactions on Pattern Analysis and Machine Intelligence
JF - IEEE Transactions on Pattern Analysis and Machine Intelligence
IS - 12
ER -