TY - JOUR
T1 - MRFMA
T2 - A hybrid paradigm integrating multi-receptive field network with mediator attention for 3D multi-organ segmentation
AU - Cui, Hengfei
AU - Li, Jiatong
AU - Du, Dianrong
AU - Zhang, Yanning
AU - Xia, Yong
N1 - Publisher Copyright:
© 2025 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/3/5
Y1 - 2026/3/5
N2 - Multi-organ segmentation has become a critical task in medical image analysis, and a precise understanding of anatomical structures is crucial for advancing disease diagnosis, treatment planning and prognosis. Existing three-dimensional (3D) multi-organ segmentation algorithms usually combine 3D Convolutional Neural Networks (CNNs) with Transformers, for the purpose of capturing local and global features. However, traditional CNNs with a fixed size of receptive field struggle to adapt to the diverse scales and long-range spatial relationships of multiple organs. Transformers have been widely used to establish dependencies on global information, despite this, they greatly increase the computational complexity. To mitigate these challenges, we propose a hybrid paradigm, called Multi-Receptive Field Network with Mediator Attention (MRFMA), to boost the representation quality for robust multi-organ segmentation across diverse imaging modalities. In MRFMA, the novel multi-receptive field depthwise convolutional module adeptly preserves the inherent inductive biases of convolution while enhancing the network’s capacity to model both fine-grained local patterns and long-range contextual relationships across anatomically disparate organs. Besides, a novel attention mechanism called Mediator Attention is developed to establish dependencies on global information. Mediator attention avoids the direct similarity calculation of query (Q) and key (K) by introducing the Mediator tokens, and thus dramatically decreases the computational cost. The proposed MRFMA is tested on the MM-WHS 2017 CT and MR dataset, FLARE 2021 dataset as well as the BTCV dataset, achieving the average Dice scores of 93.7 %, 82.2 %, 93.0 % and 83.2 % respectively. Extensive experimental results prove that our proposed method achieves superior performances in comparison with state-of-the-art (SOTA) methods. Our code will be released viahttps://github.com/jiatong0925/MRFTA.
AB - Multi-organ segmentation has become a critical task in medical image analysis, and a precise understanding of anatomical structures is crucial for advancing disease diagnosis, treatment planning and prognosis. Existing three-dimensional (3D) multi-organ segmentation algorithms usually combine 3D Convolutional Neural Networks (CNNs) with Transformers, for the purpose of capturing local and global features. However, traditional CNNs with a fixed size of receptive field struggle to adapt to the diverse scales and long-range spatial relationships of multiple organs. Transformers have been widely used to establish dependencies on global information, despite this, they greatly increase the computational complexity. To mitigate these challenges, we propose a hybrid paradigm, called Multi-Receptive Field Network with Mediator Attention (MRFMA), to boost the representation quality for robust multi-organ segmentation across diverse imaging modalities. In MRFMA, the novel multi-receptive field depthwise convolutional module adeptly preserves the inherent inductive biases of convolution while enhancing the network’s capacity to model both fine-grained local patterns and long-range contextual relationships across anatomically disparate organs. Besides, a novel attention mechanism called Mediator Attention is developed to establish dependencies on global information. Mediator attention avoids the direct similarity calculation of query (Q) and key (K) by introducing the Mediator tokens, and thus dramatically decreases the computational cost. The proposed MRFMA is tested on the MM-WHS 2017 CT and MR dataset, FLARE 2021 dataset as well as the BTCV dataset, achieving the average Dice scores of 93.7 %, 82.2 %, 93.0 % and 83.2 % respectively. Extensive experimental results prove that our proposed method achieves superior performances in comparison with state-of-the-art (SOTA) methods. Our code will be released viahttps://github.com/jiatong0925/MRFTA.
KW - Mediator attention
KW - Multi-modality
KW - Multi-organ segmentation
KW - Multi-receptive field
KW - Transformer
UR - https://www.scopus.com/pages/publications/105024356665
U2 - 10.1016/j.eswa.2025.130447
DO - 10.1016/j.eswa.2025.130447
M3 - 文章
AN - SCOPUS:105024356665
SN - 0957-4174
VL - 300
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 130447
ER -