跳到主要导航 跳到搜索 跳到主要内容

C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning

  • Northwestern Polytechnical University Xian
  • The National Engineering Laboratory for Integrated Aerospace-Ground-Ocean Big Data Application Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

2 引用 (Scopus)

摘要

Vision-Language Instruction Tuning (VLIT) is a critical training phase for Large Vision-Language Models (LVLMs). With the improving capabilities of open-source LVLMs, researchers have increasingly turned to generate VLIT data by using open-source LVLMs and achieved significant progress. However, such data generation approaches are bottlenecked by the following challenges: 1) Since multi-modal models tend to be influenced by prior language knowledge, directly using LVLMs to generate VLIT data would inevitably lead to low content relevance between generated data and images. 2) To improve the ability of the models to generate VLIT data, previous methods have incorporated an additional training phase to boost the generative capacity. This process hurts the generalization of the models to unseen inputs (i.e., “exposure bias” problem). In this paper, we propose a new Content Correlated VLIT data generation via Contrastive Learning (C3L). Specifically, we design a new content relevance module which enhances the content relevance between VLIT data and images by computing Image Instruction Correspondence Scores S(I2C). Moreover, a contrastive learning module is introduced to further boost the VLIT data generation capability of the LVLMs. A large number of automatic measures on four benchmarks show the effectiveness of our method.

源语言英语
主期刊名Proceedings of the 33rd International Joint Conference on Artificial Intelligence, IJCAI 2024
编辑Kate Larson
出版商International Joint Conferences on Artificial Intelligence
1155-1163
页数9
ISBN(电子版)9781956792041
出版状态已出版 - 2024
活动33rd International Joint Conference on Artificial Intelligence, IJCAI 2024 - Jeju, 韩国
期限: 3 8月 20249 8月 2024

丛书

姓名IJCAI International Joint Conference on Artificial Intelligence
ISSN(印刷版)1045-0823

会议

会议33rd International Joint Conference on Artificial Intelligence, IJCAI 2024
国家/地区韩国
Jeju
时期3/08/249/08/24

学术指纹

探究 'C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此