• Home
  • Search
  • Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models.
  • Cite Icon20
  • https://doi.org/10.1109/tpami.2025.3637265Copy DOI Icon

Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models.

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performance gap persists compared to advanced models like GPT-4 and Gemini. We propose a novel approach to narrow the gap by mining the potential of VLMs for better performance across various cross-modal tasks. It tackles the following questions: (1) How can high-resolution visual tokens improve image understanding without lengthening the token sequence? (2) How to improve reasoning and generation abilities of VLM with high-quality data? (3) How to close the gap between open-source VLMs and proprietary models on reasoning-driven generation? In particular, to enhance visual tokens, we propose to utilize an additional visual encoder for high-resolution refinement without increasing the visual token count. We further construct a high-quality dataset that promotes precise image comprehension and reasoning-based generation, expanding the operational scope of current VLMs. In general, Mini-Gemini further mines the potential of VLMs and empowers current frameworks with image understanding, reasoning, and generation simultaneously. The proposed model supports a series of dense and MoE Large Language Models (LLMs) from 2B to 34B, which achieve leading performance in several zero-shot benchmarks and even surpasses the developed private models. It is demonstrated to attain 80.6% accuracy on the MMB benchmark (+5.4 vs Gemini Pro) and 74.1% on TextVQA (+4.6 vs LLaVA-NeXT), achieving leading performance in several zero-shot benchmarks and even surpasses the developed private models. Furthermore, Mini-Gemini is proven to improve consistently with stronger LLM, visual encoder, and data in experiments.

Similar Papers
  • Research Article
  • Citations22

VIGC: Visual Instruction Generation and Correction

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Bin Wang +10
  • Research Article

Diagnostic accuracy of large language models in the classification of superior labial frenulum attachments.

  • Dec 10, 2025
  • Odontology
  • Mehmet Gümüş Kanmaz +1
  • Research Article
  • Citations28

LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Senqiao Yang +10
  • Supplementary Content
  • Citations24

AI-enabled language models (LMs) to large language models (LLMs) and multimodal large language models (MLLMs) in drug discovery and development

  • Feb 12, 2025
  • Journal of Advanced Research
  • Chiranjib Chakraborty +5
  • Research Article
  • Citations16

A Review of the Opportunities and Challenges with Large Language Models in Radiology: The Road Ahead.

  • Nov 21, 2024
  • AJNR. American journal of neuroradiology
  • Neetu Soni +4
  • Conference Article
  • Citations14

Evaluating Language Models for Generating and Judging Programming Feedback

  • Feb 12, 2025
  • Charles Koutcheme +6
  • Research Article

Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation

  • Nov 19, 2025
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Xiaobo Yang +1
  • Research Article
  • Citations6

A large language model for multimodal identification of crop diseases and pests

  • Jul 01, 2025
  • Scientific Reports
  • Yiqun Wang +7
  • Conference Article

Multimodal Analysis of Google Bard: Experiments in Visual Reasoning

  • Nov 25, 2023
  • David Noever +1
  • Research Article
  • Citations1

Orthodontic Biomechanical Reasoning with Multimodal Language Models: Performance and Clinical Utility

  • Oct 27, 2025
  • Bioengineering
  • Arda Arısan +2
  • Front Matter
  • Citations5

Overview of South Korean Guidelines for Approval of Large Language or Multimodal Models as Medical Devices: Key Features and Areas for Improvement.

  • Jan 01, 2025
  • Korean journal of radiology
  • Seong Ho Park +3
  • Supplementary Content

Large language and vision-language models for robot: safety challenges, mitigation strategies and future directions

  • Jul 29, 2025
  • Industrial Robot: the international journal of robotics research and application
  • Xiangyu Hu +1
  • Research Article

Influence of structured output constraints on GPT-5-Thinking, Gemini 2.5 Pro, and open-weight LLMs for radiology protocol selection.

  • Apr 10, 2026
  • European radiology experimental
  • Mohammed Bahaaeldin +9
  • Research Article
  • Citations32

Visual cognition in multimodal large language models

  • Jan 01, 2025
  • Nature Machine Intelligence
  • Luca M Schulze Buschoff +3
  • PDF
  • Conference Article
  • Citations2

Test Large Language Models on Driving Theory Knowledge and Skills for Connected Autonomous Vehicles

  • Nov 18, 2024
  • Zuoyin Tang +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.