跳到主要导航 跳到搜索 跳到主要内容

Advancing image-to-text: MDE-net coupled with large language models for context-aware captioning

  • Northwestern Polytechnical University Xian
  • South China University of Technology

科研成果: 期刊稿件文章同行评审

摘要

Image-to-text conversion is an active cross-modal research area that integrates visual and linguistic models. Currently, end-to-end image captioning models have achieved significant progress. However, conventional visual networks suffer from limited receptive fields and hierarchical feature degradation, which impede fine-grained detail extraction and discriminative recognition in complex scenes. Additionally, due to the relatively simple annotations of existing datasets, the richness and expressiveness of generated captions leave substantial room for improvement. To address these issues, this paper proposes the Multi-scale Discriminative Enhancement Network (MDE-Net), which consists of two main components. The first component is a Transformer framework integrating an Inner-embedded Multi-scale Attention (IMA) module. Unlike approaches that rely on multi-stage processes or external feature extraction networks to enrich image information, this framework fuses multi-scale features directly within the attention head space. The second component builds upon the aforementioned framework by incorporating an Adaptive Frequency Fusion (AFF) module. This module effectively leverages the complementary characteristics of detail and edge information represented across different frequencies, aiding in the precise identification of objects and their relationships. Furthermore, this work introduces a vision-guided LLM refinement module that uses language priors to improve caption fluency, lexical diversity, and descriptive expression under visual keyword constraints. On the MS COCO dataset, MDE-Net improves all standard captioning metrics over the baseline, including a 4.2-point gain in CIDEr. Further ablation and transferability experiments verify the effectiveness of the proposed modules, while the constrained LLM refinement achieves favorable performance in caption richness and expressiveness. Code is available at: https://github.com/yangyanggit89/MDE-Net.

源语言英语
文章编号133552
期刊Expert Systems with Applications
332
DOI
出版状态已出版 - 1 1月 2027

学术指纹

探究 'Advancing image-to-text: MDE-net coupled with large language models for context-aware captioning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此