Abstract
Multimedia content analysis demands accurate pixel-level semantic segmentation that tightly couples visual details with rich textual context. Yet existing methods largely neglect the dense semantic knowledge inherently embedded within visual objects, resulting in insufficient fine-grained alignment between semantic concepts and spatial visual details, thereby limiting segmentation accuracy. We propose DK-Seg (Deep Knowledge Segmentation), a deeply integrated layer-wise semantic knowledge field framework that reformulates segmentation as the alignment between a dense, attribute-rich semantic knowledge field and the visual feature hierarchy of a SAM-based encoder. Instead of attaching language only at the input or output, DK-Seg injects dense semantic knowledge throughout the encoder, enabling each layer to operate under structured semantic conditioning. Our framework comprises three theoretically grounded alignment mechanisms: (1) dynamic layer-wise cross-attention with adaptive residual gating, enabling controlled semantic modulation across the hierarchy; (2) fine-grained pixel-text alignment, establishing token-level correspondences between semantic attributes and spatial patterns;(3) global optimal transport alignment, which matches visual and textual embedding distributions under a Wasserstein-2 geometry to ensure global consistency. Extensive evaluations on high-quality segmentation, referring segmentation, and additional semantic segmentation benchmarks demonstrate that DK-Seg consistently improves over state-of-the-art models. We further include computational complexity analysis and comprehensive ablations confirming that the improvements arise from the proposed modeling framework rather than from architectural scale alone.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Dense Knowledge Infusion
- Layer-wise Cross-Attention
- Multimodal Fusion
- Multitimedia
- Semantic Segmentation
Fingerprint
Dive into the research topics of 'Deep Layer-wise Infusion and Alignment of Dense Knowledge for Semantic Segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver