Abstract
Image-to-text conversion is an active cross-modal research area that integrates visual and linguistic models. Currently, end-to-end image captioning models have achieved significant progress. However, conventional visual networks suffer from limited receptive fields and hierarchical feature degradation, which impede fine-grained detail extraction and discriminative recognition in complex scenes. Additionally, due to the relatively simple annotations of existing datasets, the richness and expressiveness of generated captions leave substantial room for improvement. To address these issues, this paper proposes the Multi-scale Discriminative Enhancement Network (MDE-Net), which consists of two main components. The first component is a Transformer framework integrating an Inner-embedded Multi-scale Attention (IMA) module. Unlike approaches that rely on multi-stage processes or external feature extraction networks to enrich image information, this framework fuses multi-scale features directly within the attention head space. The second component builds upon the aforementioned framework by incorporating an Adaptive Frequency Fusion (AFF) module. This module effectively leverages the complementary characteristics of detail and edge information represented across different frequencies, aiding in the precise identification of objects and their relationships. Furthermore, this work introduces a vision-guided LLM refinement module that uses language priors to improve caption fluency, lexical diversity, and descriptive expression under visual keyword constraints. On the MS COCO dataset, MDE-Net improves all standard captioning metrics over the baseline, including a 4.2-point gain in CIDEr. Further ablation and transferability experiments verify the effectiveness of the proposed modules, while the constrained LLM refinement achieves favorable performance in caption richness and expressiveness. Code is available at: https://github.com/yangyanggit89/MDE-Net.
| Original language | English |
|---|---|
| Article number | 133552 |
| Journal | Expert Systems with Applications |
| Volume | 332 |
| DOIs | |
| State | Published - 1 Jan 2027 |
Keywords
- Cross-modal research
- Frequency decomposition
- Inner-embedded multi-scale attention
- LLMs
- Text generation
Fingerprint
Dive into the research topics of 'Advancing image-to-text: MDE-net coupled with large language models for context-aware captioning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver