TY - JOUR
T1 - A review of ultrasound video segmentation
T2 - From temporal modeling to clinical utility
AU - Huang, Qinghua
AU - An, Zhaoxing
AU - Li, Guangju
N1 - Publisher Copyright:
© 2026
PY - 2026/10/28
Y1 - 2026/10/28
N2 - Ultrasound is an inherently dynamic imaging modality, and in many clinical scenarios it is acquired as continuous video rather than isolated static frames. This characteristic makes ultrasound video segmentation fundamentally different from conventional image segmentation, as clinically useful analysis requires both accurate frame-level delineation and reliable temporal consistency across sequences. However, the task remains challenging due to speckle noise, low contrast, non-rigid tissue motion, probe-induced artifacts, and limited densely annotated video data. To address these issues, a broad range of methods has been developed, with particular emphasis on temporal modeling strategies that exploit inter-frame dependencies to improve segmentation robustness and consistency. Despite notable progress, existing studies remain scattered across anatomical targets, datasets, annotation protocols, evaluation criteria, and clinical tasks, hindering systematic comparison and limiting the assessment of clinical utility. This paper presents a comprehensive review of ultrasound video segmentation from the perspective of the progression from temporal modeling to clinical utility. We summarize the distinctive characteristics of this task in real-world clinical workflows, review major methodological paradigms, publicly available datasets, and commonly used evaluation strategies, and further discuss representative task-oriented clinical scenarios. Finally, we highlight current limitations and future directions toward more robust, interpretable, and clinically useful ultrasound video segmentation systems.
AB - Ultrasound is an inherently dynamic imaging modality, and in many clinical scenarios it is acquired as continuous video rather than isolated static frames. This characteristic makes ultrasound video segmentation fundamentally different from conventional image segmentation, as clinically useful analysis requires both accurate frame-level delineation and reliable temporal consistency across sequences. However, the task remains challenging due to speckle noise, low contrast, non-rigid tissue motion, probe-induced artifacts, and limited densely annotated video data. To address these issues, a broad range of methods has been developed, with particular emphasis on temporal modeling strategies that exploit inter-frame dependencies to improve segmentation robustness and consistency. Despite notable progress, existing studies remain scattered across anatomical targets, datasets, annotation protocols, evaluation criteria, and clinical tasks, hindering systematic comparison and limiting the assessment of clinical utility. This paper presents a comprehensive review of ultrasound video segmentation from the perspective of the progression from temporal modeling to clinical utility. We summarize the distinctive characteristics of this task in real-world clinical workflows, review major methodological paradigms, publicly available datasets, and commonly used evaluation strategies, and further discuss representative task-oriented clinical scenarios. Finally, we highlight current limitations and future directions toward more robust, interpretable, and clinically useful ultrasound video segmentation systems.
KW - Clinical utility
KW - Medical image analysis
KW - Spatio-temporal learning
KW - Temporal modeling
KW - Ultrasound video segmentation
UR - https://www.scopus.com/pages/publications/105043817333
U2 - 10.1016/j.neucom.2026.134407
DO - 10.1016/j.neucom.2026.134407
M3 - 文献综述
AN - SCOPUS:105043817333
SN - 0925-2312
VL - 699
JO - Neurocomputing
JF - Neurocomputing
M1 - 134407
ER -