Skip to main navigation Skip to search Skip to main content

Deep Layer-wise Infusion and Alignment of Dense Knowledge for Semantic Segmentation

  • Northwestern Polytechnical University Xian

Research output: Contribution to journalArticlepeer-review

Abstract

Multimedia content analysis demands accurate pixel-level semantic segmentation that tightly couples visual details with rich textual context. Yet existing methods largely neglect the dense semantic knowledge inherently embedded within visual objects, resulting in insufficient fine-grained alignment between semantic concepts and spatial visual details, thereby limiting segmentation accuracy. We propose DK-Seg (Deep Knowledge Segmentation), a deeply integrated layer-wise semantic knowledge field framework that reformulates segmentation as the alignment between a dense, attribute-rich semantic knowledge field and the visual feature hierarchy of a SAM-based encoder. Instead of attaching language only at the input or output, DK-Seg injects dense semantic knowledge throughout the encoder, enabling each layer to operate under structured semantic conditioning. Our framework comprises three theoretically grounded alignment mechanisms: (1) dynamic layer-wise cross-attention with adaptive residual gating, enabling controlled semantic modulation across the hierarchy; (2) fine-grained pixel-text alignment, establishing token-level correspondences between semantic attributes and spatial patterns;(3) global optimal transport alignment, which matches visual and textual embedding distributions under a Wasserstein-2 geometry to ensure global consistency. Extensive evaluations on high-quality segmentation, referring segmentation, and additional semantic segmentation benchmarks demonstrate that DK-Seg consistently improves over state-of-the-art models. We further include computational complexity analysis and comprehensive ablations confirming that the improvements arise from the proposed modeling framework rather than from architectural scale alone.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026

Keywords

  • Dense Knowledge Infusion
  • Layer-wise Cross-Attention
  • Multimodal Fusion
  • Multitimedia
  • Semantic Segmentation

Fingerprint

Dive into the research topics of 'Deep Layer-wise Infusion and Alignment of Dense Knowledge for Semantic Segmentation'. Together they form a unique fingerprint.

Cite this