Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction

Jingfen Qiao; Jia-Huei Ju; Xinyu Ma; Evangelos Kanoulas; Andrew Yates

doi:https://doi.org/10.1145/3726302.3730285

Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction

Authors	Jingfen Qiao Jia-Huei Ju Xinyu Ma Evangelos Kanoulas Andrew Yates
Publication date	2025
Book title	SIGIR '25
Book subtitle	Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval : July 13-18, 2025, Padua, Italy
ISBN (electronic)	9798400715921
Event	48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025
Pages (from-to)	3335-3345
Number of pages	11
Publisher	New York, NY: Association for Computing Machinery
Organisations	Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract	Visual Document Retrieval (VDR) is an emerging research area that focuses on encoding and retrieving document images directly, bypassing the dependence on Optical Character Recognition (OCR) for document search. A recent advance in VDR was introduced by ColPali, which significantly improved retrieval effectiveness through a late interaction mechanism. ColPali’s approach demonstrated substantial performance gains over existing baselines that do not use late interaction on an established benchmark. In this study, we investigate the reproducibility and replicability of VDR methods with and without late interaction mechanisms by systematically evaluating their performance across multiple pre-trained vision-language models. Our findings confirm that late interaction yields considerable improvements in retrieval effectiveness; however, it also introduces computational inefficiencies during inference. Additionally, we examine the adaptability of VDR models to textual inputs and assess their robustness across text-intensive datasets within the proposed benchmark, particularly when scaling the indexing mechanism. Furthermore, our research investigates the specific contributions of late interaction by looking into query-patch matching in the context of visual document retrieval. We find that although query tokens cannot explicitly match image patches as in the text retrieval scenario, they tend to match the patch contains visually similar tokens or their surrounding patches.
Document type	Conference contribution
Language	English
Published at	https://doi.org/10.1145/3726302.3730285 (Final published version)
Other links	https://www.scopus.com/pages/publications/105011828552
Downloads	3726302.3730285 (Final published version)
Permalink to this page

Back

UvA-DARE

Digital Academic Repository

Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction