Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

Open Access
Authors
Publication date 2025
Book title 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition : CVPR 2025
Book subtitle Nashville, Tennessee, USA, 11-15 June 2025 : proceedings
ISBN
  • 9798331543655
ISBN (electronic)
  • 9798331543648
Event 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025
Pages (from-to) 3262-3271
Publisher Los Alamitos, California: IEEE Computer Society
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing concerns that they learn spurious correlations between videos and likely answers, reinforced by their black-box nature and remaining benchmarking biases. Our method explicitly grounds VQA tasks to video fragments in four steps: entailment tree construction, video-language entailment verification, tree reasoning, and dynamic tree expansion. A vital benefit of the method is its generalizability to current video- and image-based VLMs across reasoning types. To support fair evaluation, we devise a de-biasing procedure based on large-language models that rewrites VQA benchmark answer sets to enforce model reasoning. Systematic experiments on existing and de-biased benchmarks highlight the impact of our method components across benchmarks, VLMs, and reasoning types.
Document type Conference contribution
Note With supplemental file
Language English
Published at
https://doi.org/10.48550/arXiv.2501.05069 (Accepted author manuscript)
Published at
Other links
Downloads
Supplementary materials
Permalink to this page
Back