- Research Article
- 10.3390/electronics15071353
Accurate Multi-Page Document Retrieval by Effectively Fusing Context Information Across Pages
- Mar 25, 2026
- Electronics
- Bing Qian + 6 more +6
Visual retrievers such as visual retrieval-augmented generation (RAG) have recently emerged as a powerful model for retrieving multimodal documents without the need to convert page images into text. Existing visual retrievers typically encode every document page separately, ignoring the inherent rich context information across pages within multi-page documents. However, some crucial semantic information often spans multiple pages in a document, and should be effectively encoded for better retrieval. To address this problem, this paper proposes a novel approach utilizing dynamically fusing visual context (DFVC), which adaptively encodes the semantic information across pages. In the proposed DFVC approach, a lightweight plug-and-play adapter is designed; in addition, a contrastive loss function incorporating the positive fused embedding vectors and negative embedding vectors is designed to constrain the adapter, allowing it to learn the weights for the context pages. Together, the designed adapter and loss function allow the retriever to effectively encode useful semantic information across pages while excluding distracting noise. The proposed DFVC is validated on commonly used challenging multi-page document benchmarks. Extensive experimental results demonstrate that it significantly boosts retrieval performance. In addition, the proposed DFVC is highly parameter-efficient since it employs frozen vision-language backbones, allowing it to be easily integrated into existing visual RAG pipelines for finer document retrieval.
Read more