• https://doi.org/10.1109/icip55913.2025.11084631Copy DOI Icon

Visual Prompting Through Image Mines

  • Sep 14, 2025
  • Kalash Shah +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Visual prompting aims to enhance the performance of vision-language models (VLMs), which, despite their remarkable capabilities, often struggle with dense, detailed images, leading to incorrect answers or hallucinations. We propose Visual Prompting Through Image Mines , a novel algorithm that leverages attention patterns from a base VLM to generate image crops for improved visual grounding. Specifically, we extract attention values from output text tokens in the LLaVA-8B model, overlay them onto image patches to create an attention graph, and apply a modified breadth-first search (BFS) to identify key regions as image crops. Using the SigLIP model, we refine these regions into Image Mines, retaining only the most relevant crops. Our approach supports both single-image and multi-image inference setups, delivering superior performance. It consistently outperforms existing visual prompting methods in nearly all scenarios.

Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.