TY - JOUR
T1 - Advancing image-to-text
T2 - MDE-net coupled with large language models for context-aware captioning
AU - Yang, Yang
AU - Yu, Dengxiu
AU - Chen, C. L.Philip
N1 - Publisher Copyright:
© 2026 Published by Elsevier Ltd.
PY - 2027/1/1
Y1 - 2027/1/1
N2 - Image-to-text conversion is an active cross-modal research area that integrates visual and linguistic models. Currently, end-to-end image captioning models have achieved significant progress. However, conventional visual networks suffer from limited receptive fields and hierarchical feature degradation, which impede fine-grained detail extraction and discriminative recognition in complex scenes. Additionally, due to the relatively simple annotations of existing datasets, the richness and expressiveness of generated captions leave substantial room for improvement. To address these issues, this paper proposes the Multi-scale Discriminative Enhancement Network (MDE-Net), which consists of two main components. The first component is a Transformer framework integrating an Inner-embedded Multi-scale Attention (IMA) module. Unlike approaches that rely on multi-stage processes or external feature extraction networks to enrich image information, this framework fuses multi-scale features directly within the attention head space. The second component builds upon the aforementioned framework by incorporating an Adaptive Frequency Fusion (AFF) module. This module effectively leverages the complementary characteristics of detail and edge information represented across different frequencies, aiding in the precise identification of objects and their relationships. Furthermore, this work introduces a vision-guided LLM refinement module that uses language priors to improve caption fluency, lexical diversity, and descriptive expression under visual keyword constraints. On the MS COCO dataset, MDE-Net improves all standard captioning metrics over the baseline, including a 4.2-point gain in CIDEr. Further ablation and transferability experiments verify the effectiveness of the proposed modules, while the constrained LLM refinement achieves favorable performance in caption richness and expressiveness. Code is available at: https://github.com/yangyanggit89/MDE-Net.
AB - Image-to-text conversion is an active cross-modal research area that integrates visual and linguistic models. Currently, end-to-end image captioning models have achieved significant progress. However, conventional visual networks suffer from limited receptive fields and hierarchical feature degradation, which impede fine-grained detail extraction and discriminative recognition in complex scenes. Additionally, due to the relatively simple annotations of existing datasets, the richness and expressiveness of generated captions leave substantial room for improvement. To address these issues, this paper proposes the Multi-scale Discriminative Enhancement Network (MDE-Net), which consists of two main components. The first component is a Transformer framework integrating an Inner-embedded Multi-scale Attention (IMA) module. Unlike approaches that rely on multi-stage processes or external feature extraction networks to enrich image information, this framework fuses multi-scale features directly within the attention head space. The second component builds upon the aforementioned framework by incorporating an Adaptive Frequency Fusion (AFF) module. This module effectively leverages the complementary characteristics of detail and edge information represented across different frequencies, aiding in the precise identification of objects and their relationships. Furthermore, this work introduces a vision-guided LLM refinement module that uses language priors to improve caption fluency, lexical diversity, and descriptive expression under visual keyword constraints. On the MS COCO dataset, MDE-Net improves all standard captioning metrics over the baseline, including a 4.2-point gain in CIDEr. Further ablation and transferability experiments verify the effectiveness of the proposed modules, while the constrained LLM refinement achieves favorable performance in caption richness and expressiveness. Code is available at: https://github.com/yangyanggit89/MDE-Net.
KW - Cross-modal research
KW - Frequency decomposition
KW - Inner-embedded multi-scale attention
KW - LLMs
KW - Text generation
UR - https://www.scopus.com/pages/publications/105044386799
U2 - 10.1016/j.eswa.2026.133552
DO - 10.1016/j.eswa.2026.133552
M3 - 文章
AN - SCOPUS:105044386799
SN - 0957-4174
VL - 332
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 133552
ER -