• Home
  • Search
  • VIGC: Visual Instruction Generation and Correction
  • Cite Icon22
  • https://doi.org/10.1609/aaai.v38i6.28338Copy DOI Icon

VIGC: Visual Instruction Generation and Correction

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality instruction-tuning data for vision-language tasks remains a challenge. The current leading paradigm, such as LLaVA, relies on language-only GPT-4 to generate data, which requires pre-annotated image captions and detection bounding boxes, suffering from understanding image details. A practical solution to this problem would be to utilize the available multimodal large language models to generate instruction data for vision-language tasks. However, it's worth noting that the currently accessible MLLMs are not as powerful as their LLM counterparts, as they tend to produce inadequate responses and generate false information. As a solution for addressing the current issue, this paper proposes the Visual Instruction Generation and Correction (VIGC) framework that enables multimodal large language models to generate instruction-tuning data and progressively enhance its quality on-the-fly. Specifically, Visual Instruction Generation (VIG) guides the vision-language model to generate diverse instruction-tuning data. To ensure generation quality, Visual Instruction Correction (VIC) adopts an iterative update mechanism to correct any inaccuracies in data produced by VIG, effectively reducing the risk of hallucination. Leveraging the diverse, high-quality data generated by VIGC, we finetune mainstream models and validate data quality based on various evaluations. Experimental results demonstrate that VIGC not only compensates for the shortcomings of language-only data generation methods, but also effectively enhances the benchmark performance. The models, datasets, and code are available at https://opendatalab.github.io/VIGC

Similar Papers
  • Research Article

Diagnostic accuracy of large language models in the classification of superior labial frenulum attachments.

  • Dec 10, 2025
  • Odontology
  • Mehmet Gümüş Kanmaz +1
  • Research Article
  • Citations20

Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models.

  • Mar 01, 2026
  • IEEE transactions on pattern analysis and machine intelligence
  • Yanwei Li +7
  • Research Article
  • Citations28

LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Senqiao Yang +10
  • Research Article
  • Citations14

Evaluating the strengths and limitations of multimodal ChatGPT-4 in detecting glaucoma using fundus images.

  • Jun 07, 2024
  • Frontiers in ophthalmology
  • Saif Aldeen Alryalat +2
  • Conference Article

Enhancing Sentiment Analysis with Multimodal Large Language Models

  • May 23, 2025
  • Thresa Jeniffer J +4
  • Supplementary Content
  • Citations24

AI-enabled language models (LMs) to large language models (LLMs) and multimodal large language models (MLLMs) in drug discovery and development

  • Feb 12, 2025
  • Journal of Advanced Research
  • Chiranjib Chakraborty +5
  • Research Article

Large-scale evaluation of multimodal large language models for pneumothorax detection.

  • Feb 12, 2026
  • Diagnostic and interventional radiology (Ankara, Turkey)
  • Hamza Eren Güzel +2
  • Research Article

The Power of Multimodality in Multimodal Large Language Models, Unimodal ChatGPT 5.0, and Human Clinical Experts on a Wound Care Certification Examination: Cross-Sectional Comparative Study.

  • Apr 27, 2026
  • JMIR formative research
  • Mete Ucdal +5
  • Research Article
  • Citations8

Medical MLLM Is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Xijie Huang +8
  • Conference Article

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

  • Dec 08, 2025
  • Kimia Ehsani +1
  • Research Article

TinnitusLLM: A Multimodal Large Language Model Framework for Tinnitus Diagnosis Through EEG-fMRI Fusion Learning.

  • Jan 01, 2026
  • IEEE journal of biomedical and health informatics
  • Yipeng Du +9
  • Research Article
  • Citations15

Embodied Intelligence in Mining: Leveraging Multi-Modal Large Language Models for Autonomous Driving in Mines

  • May 01, 2024
  • IEEE Transactions on Intelligent Vehicles
  • Luxi Li +9
  • Research Article
  • Citations16

From Decision to Action in Surgical Autonomy: Multi-Modal Large Language Models for Robot-Assisted Blood Suction

  • Mar 01, 2025
  • IEEE Robotics and Automation Letters
  • Sadra Zargarzadeh +3
  • Research Article

Evaluating multimodal commercial and open-source large language models for dynamical astronomy: a benchmark study of resonant behavior classification.

  • Mar 28, 2026
  • Scientific reports
  • Evgeny Smirnov +1
  • Supplementary Content
  • Citations3

Multimodal large language models for oral lesion diagnosis: a systematic review of diagnostic performance and clinical utility

  • Feb 24, 2026
  • Frontiers in Oral Health
  • Fatma E A Hassanein +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.