Skip to main navigation Skip to search Skip to main content

Advancing image-to-text: MDE-net coupled with large language models for context-aware captioning

  • Northwestern Polytechnical University Xian
  • South China University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Image-to-text conversion is an active cross-modal research area that integrates visual and linguistic models. Currently, end-to-end image captioning models have achieved significant progress. However, conventional visual networks suffer from limited receptive fields and hierarchical feature degradation, which impede fine-grained detail extraction and discriminative recognition in complex scenes. Additionally, due to the relatively simple annotations of existing datasets, the richness and expressiveness of generated captions leave substantial room for improvement. To address these issues, this paper proposes the Multi-scale Discriminative Enhancement Network (MDE-Net), which consists of two main components. The first component is a Transformer framework integrating an Inner-embedded Multi-scale Attention (IMA) module. Unlike approaches that rely on multi-stage processes or external feature extraction networks to enrich image information, this framework fuses multi-scale features directly within the attention head space. The second component builds upon the aforementioned framework by incorporating an Adaptive Frequency Fusion (AFF) module. This module effectively leverages the complementary characteristics of detail and edge information represented across different frequencies, aiding in the precise identification of objects and their relationships. Furthermore, this work introduces a vision-guided LLM refinement module that uses language priors to improve caption fluency, lexical diversity, and descriptive expression under visual keyword constraints. On the MS COCO dataset, MDE-Net improves all standard captioning metrics over the baseline, including a 4.2-point gain in CIDEr. Further ablation and transferability experiments verify the effectiveness of the proposed modules, while the constrained LLM refinement achieves favorable performance in caption richness and expressiveness. Code is available at: https://github.com/yangyanggit89/MDE-Net.

Original languageEnglish
Article number133552
JournalExpert Systems with Applications
Volume332
DOIs
StatePublished - 1 Jan 2027

Keywords

  • Cross-modal research
  • Frequency decomposition
  • Inner-embedded multi-scale attention
  • LLMs
  • Text generation

Fingerprint

Dive into the research topics of 'Advancing image-to-text: MDE-net coupled with large language models for context-aware captioning'. Together they form a unique fingerprint.

Cite this