TY - GEN
T1 - Hierarchical Feature Integration Network for RGB-D Saliency Detection
AU - Guo, Kuo
AU - Guo, Yangming
AU - Yao, Wengang
AU - Fan, Yongxin
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Despite their promising results, existing two-stream RGB-D saliency detection methods fall short in fully exploiting cross-modal complementary features. In this paper, we develop a novel hierarchical feature integration network (HFINet), designed to explicitly and effectively leverage the complementary nature of two-stream features while integrating multi-level spatial characteristics. Specifically, for low-level features where RGB contains richer detail than depth, we introduce a pyramid spatial fusion module (PSFM) to extract only RGB-based detail information, enhancing both detail preservation and contextual transmission. For high-level features, a pyramid feature fusion module (PFFM) is proposed to capture semantic content from RGB while aggregating contextual fusion information. Moreover, a feature interaction module (FIM) is designed to leverage depth cues to assist RGB representation, enabling accurate mining of semantic complementarity between the two modalities. Finally, a lightweight feature fusion decoder is adopted to facilitate efficient feature transformation from the encoder to the decoder. Extensive experiments on several datasets show that HFINet achieves competitive performance compared to 11 representative methods.
AB - Despite their promising results, existing two-stream RGB-D saliency detection methods fall short in fully exploiting cross-modal complementary features. In this paper, we develop a novel hierarchical feature integration network (HFINet), designed to explicitly and effectively leverage the complementary nature of two-stream features while integrating multi-level spatial characteristics. Specifically, for low-level features where RGB contains richer detail than depth, we introduce a pyramid spatial fusion module (PSFM) to extract only RGB-based detail information, enhancing both detail preservation and contextual transmission. For high-level features, a pyramid feature fusion module (PFFM) is proposed to capture semantic content from RGB while aggregating contextual fusion information. Moreover, a feature interaction module (FIM) is designed to leverage depth cues to assist RGB representation, enabling accurate mining of semantic complementarity between the two modalities. Finally, a lightweight feature fusion decoder is adopted to facilitate efficient feature transformation from the encoder to the decoder. Extensive experiments on several datasets show that HFINet achieves competitive performance compared to 11 representative methods.
KW - Salient object detection
KW - interactive attention
KW - pyramid spatial fusion
KW - semantic complementarity
UR - https://www.scopus.com/pages/publications/105042565977
U2 - 10.1007/978-981-95-8232-7_30
DO - 10.1007/978-981-95-8232-7_30
M3 - 会议稿件
AN - SCOPUS:105042565977
SN - 9789819582310
T3 - Lecture Notes in Electrical Engineering
SP - 311
EP - 326
BT - Proceedings of the 4th International Conference on Sensing, Measurement, Communication and Internet of Things Technologies
A2 - Zhao, Zhenyu
A2 - Jin, Peiquan
A2 - Zhang, Mingchuan
PB - Springer Science and Business Media Deutschland GmbH
T2 - 4th International Conference on Sensing, Measurement, Communication and Internet of Things Technologies, SMC-IoT 2025
Y2 - 28 November 2025 through 30 November 2025
ER -