TY - JOUR
T1 - Deep Layer-wise Infusion and Alignment of Dense Knowledge for Semantic Segmentation
AU - Huang, Dengdian
AU - Zhao, Bin
AU - Yuan, Yuan
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Multimedia content analysis demands accurate pixel-level semantic segmentation that tightly couples visual details with rich textual context. Yet existing methods largely neglect the dense semantic knowledge inherently embedded within visual objects, resulting in insufficient fine-grained alignment between semantic concepts and spatial visual details, thereby limiting segmentation accuracy. We propose DK-Seg (Deep Knowledge Segmentation), a deeply integrated layer-wise semantic knowledge field framework that reformulates segmentation as the alignment between a dense, attribute-rich semantic knowledge field and the visual feature hierarchy of a SAM-based encoder. Instead of attaching language only at the input or output, DK-Seg injects dense semantic knowledge throughout the encoder, enabling each layer to operate under structured semantic conditioning. Our framework comprises three theoretically grounded alignment mechanisms: (1) dynamic layer-wise cross-attention with adaptive residual gating, enabling controlled semantic modulation across the hierarchy; (2) fine-grained pixel-text alignment, establishing token-level correspondences between semantic attributes and spatial patterns;(3) global optimal transport alignment, which matches visual and textual embedding distributions under a Wasserstein-2 geometry to ensure global consistency. Extensive evaluations on high-quality segmentation, referring segmentation, and additional semantic segmentation benchmarks demonstrate that DK-Seg consistently improves over state-of-the-art models. We further include computational complexity analysis and comprehensive ablations confirming that the improvements arise from the proposed modeling framework rather than from architectural scale alone.
AB - Multimedia content analysis demands accurate pixel-level semantic segmentation that tightly couples visual details with rich textual context. Yet existing methods largely neglect the dense semantic knowledge inherently embedded within visual objects, resulting in insufficient fine-grained alignment between semantic concepts and spatial visual details, thereby limiting segmentation accuracy. We propose DK-Seg (Deep Knowledge Segmentation), a deeply integrated layer-wise semantic knowledge field framework that reformulates segmentation as the alignment between a dense, attribute-rich semantic knowledge field and the visual feature hierarchy of a SAM-based encoder. Instead of attaching language only at the input or output, DK-Seg injects dense semantic knowledge throughout the encoder, enabling each layer to operate under structured semantic conditioning. Our framework comprises three theoretically grounded alignment mechanisms: (1) dynamic layer-wise cross-attention with adaptive residual gating, enabling controlled semantic modulation across the hierarchy; (2) fine-grained pixel-text alignment, establishing token-level correspondences between semantic attributes and spatial patterns;(3) global optimal transport alignment, which matches visual and textual embedding distributions under a Wasserstein-2 geometry to ensure global consistency. Extensive evaluations on high-quality segmentation, referring segmentation, and additional semantic segmentation benchmarks demonstrate that DK-Seg consistently improves over state-of-the-art models. We further include computational complexity analysis and comprehensive ablations confirming that the improvements arise from the proposed modeling framework rather than from architectural scale alone.
KW - Dense Knowledge Infusion
KW - Layer-wise Cross-Attention
KW - Multimodal Fusion
KW - Multitimedia
KW - Semantic Segmentation
UR - https://www.scopus.com/pages/publications/105045324456
U2 - 10.1109/TMM.2026.3713784
DO - 10.1109/TMM.2026.3713784
M3 - 文章
AN - SCOPUS:105045324456
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -