TY - JOUR
T1 - Generative morphodynamic forecasting enables robust zero-shot volumetric medical segmentation
AU - Dai, Duwei
AU - Dong, Caixia
AU - Dai, Guowei
AU - Li, Xiaoli
AU - Liu, Fan
AU - Yang, Xu
AU - Qin, Bowen
AU - Yan, Qingsen
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/9
Y1 - 2026/9
N2 - Video foundation models show strong potential for interactive volumetric medical image parsing but can suffer from memory drift when applied directly to 3D medical volumes with severe, patient-specific topological changes. In the evaluated settings, standard reactive memory architectures and geometric prompts provide limited guidance for newly appearing, fragmented, or low-contrast structures, which can lead to accumulated segmentation errors. To mitigate this issue, we introduce a generative anticipatory framework that augments reactive tracking with morphological forecasting. By combining the semantic reasoning of large language models with the continuous latent dynamics of neural stochastic differential equations, our approach forecasts plausible anatomical trajectories. We propose an uncertainty-guided dual-stream memory architecture to integrate these forecasted priors with noisy visual evidence. This mechanism uses predictive spatial variance to balance empirical visual observations and forecasted priors, improving topological consistency under challenging imaging conditions in our benchmarks. We further formulate human-in-the-loop interactive corrections via Bayesian state resetting, translating expert interventions into more efficient volumetric correction. Evaluations across diverse clinical datasets show strong training-free generalization, improved robustness to morphodynamic variation, and higher interactive efficiency among the evaluated methods. These results suggest a promising uncertainty-aware direction for promptable volumetric medical segmentation.
AB - Video foundation models show strong potential for interactive volumetric medical image parsing but can suffer from memory drift when applied directly to 3D medical volumes with severe, patient-specific topological changes. In the evaluated settings, standard reactive memory architectures and geometric prompts provide limited guidance for newly appearing, fragmented, or low-contrast structures, which can lead to accumulated segmentation errors. To mitigate this issue, we introduce a generative anticipatory framework that augments reactive tracking with morphological forecasting. By combining the semantic reasoning of large language models with the continuous latent dynamics of neural stochastic differential equations, our approach forecasts plausible anatomical trajectories. We propose an uncertainty-guided dual-stream memory architecture to integrate these forecasted priors with noisy visual evidence. This mechanism uses predictive spatial variance to balance empirical visual observations and forecasted priors, improving topological consistency under challenging imaging conditions in our benchmarks. We further formulate human-in-the-loop interactive corrections via Bayesian state resetting, translating expert interventions into more efficient volumetric correction. Evaluations across diverse clinical datasets show strong training-free generalization, improved robustness to morphodynamic variation, and higher interactive efficiency among the evaluated methods. These results suggest a promising uncertainty-aware direction for promptable volumetric medical segmentation.
KW - 3D medical image segmentation
KW - Generative visual priors
KW - Neural stochastic differential
KW - Uncertainty estimation
UR - https://www.scopus.com/pages/publications/105043359436
U2 - 10.1016/j.media.2026.104180
DO - 10.1016/j.media.2026.104180
M3 - 文章
AN - SCOPUS:105043359436
SN - 1361-8415
VL - 113
JO - Medical Image Analysis
JF - Medical Image Analysis
M1 - 104180
ER -