TY - JOUR
T1 - FSCFNet
T2 - Lightweight neural networks via multi-dimensional importance-aware optimization
AU - Nie, Mengyang
AU - Sun, Jinqiu
AU - Guoyang, Hongsong
AU - Niu, Axi
AU - Hu, Yaoqi
AU - Yan, Qingsen
AU - Zhu, Yu
AU - Zhang, Yanning
N1 - Publisher Copyright:
© 2025 Elsevier B.V.
PY - 2026/1/7
Y1 - 2026/1/7
N2 - Network lightweighting has become an effective technique for compressing CNNs by eliminating redundant structures. However, several challenges remain unresolved. For instance, some convolution optimization methods decompose convolutions into multiple segments, which reduce FLOPs but simultaneously increase memory access overhead. Similarly, ensuring that quantization and pruning techniques achieve substantial improvements in computational efficiency without compromising accuracy remains a pressing challenge. Most importantly, many methods neglect hardware adaptation, resulting in no significant performance improvement. To address these limitations, we propose a framework that integrates multi-dimensional evaluation with hardware-aware optimization. 1) We introduce a multi-dimensionally important convolution module. By selectively processing only the most informative features, this module substantially reduces both Floating Point Operations per Second (FLOPs) and memory access, while maintaining accuracy. 2) We propose a mixed-precision quantization module based on multi-dimensional importance and hardware awareness. This module assigns different precision levels (high, medium, or low) to features according to their importance and how well the hardware adapts to various computational precisions. This design markedly reduces the parameter count with only marginal accuracy degradation. 3) We introduce a channel pruning module based on channel contribution assessment. Through structured pruning of channels with minimal contribution to prediction accuracy, this module further improves computational efficiency with negligible accuracy loss. We validate the proposed framework on both GPU- and CPU-based platforms. Extensive experiments demonstrate that our method improves processing speed by 2.6 times in terms of FPS and reduces the total parameter count by 71 %, all without compromising model accuracy.
AB - Network lightweighting has become an effective technique for compressing CNNs by eliminating redundant structures. However, several challenges remain unresolved. For instance, some convolution optimization methods decompose convolutions into multiple segments, which reduce FLOPs but simultaneously increase memory access overhead. Similarly, ensuring that quantization and pruning techniques achieve substantial improvements in computational efficiency without compromising accuracy remains a pressing challenge. Most importantly, many methods neglect hardware adaptation, resulting in no significant performance improvement. To address these limitations, we propose a framework that integrates multi-dimensional evaluation with hardware-aware optimization. 1) We introduce a multi-dimensionally important convolution module. By selectively processing only the most informative features, this module substantially reduces both Floating Point Operations per Second (FLOPs) and memory access, while maintaining accuracy. 2) We propose a mixed-precision quantization module based on multi-dimensional importance and hardware awareness. This module assigns different precision levels (high, medium, or low) to features according to their importance and how well the hardware adapts to various computational precisions. This design markedly reduces the parameter count with only marginal accuracy degradation. 3) We introduce a channel pruning module based on channel contribution assessment. Through structured pruning of channels with minimal contribution to prediction accuracy, this module further improves computational efficiency with negligible accuracy loss. We validate the proposed framework on both GPU- and CPU-based platforms. Extensive experiments demonstrate that our method improves processing speed by 2.6 times in terms of FPS and reduces the total parameter count by 71 %, all without compromising model accuracy.
KW - Algorithm-hardware co-optimization
KW - Feature importance assessment
KW - Hardware-aware lightweight network model
UR - https://www.scopus.com/pages/publications/105019641448
U2 - 10.1016/j.neucom.2025.131823
DO - 10.1016/j.neucom.2025.131823
M3 - 文章
AN - SCOPUS:105019641448
SN - 0925-2312
VL - 660
JO - Neurocomputing
JF - Neurocomputing
M1 - 131823
ER -