Score: 0

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Published: March 24, 2025 | arXiv ID: 2503.18923v2

By: Meng Cao , Pengfei Hu , Yingyao Wang and more

Potential Business Impact:

Tests if AI understands video facts correctly.

Business Areas:

Video Editing Content and Publishing, Media and Entertainment, Video

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailored for factuality evaluation in video contexts. Our work differs from existing video benchmarks through the following key features: 1) Knowledge required: demanding integration of external knowledge beyond the video's explicit narrative; 2) Multi-hop fact-seeking question: Each question involves multiple explicit facts and requires strict factual grounding without hypothetical or subjective inferences. We also include per-hop single-fact-based sub-QAs alongside final QAs to enable fine-grained, stepby-step evaluation; 3) Short-form definitive answer: Answers are crafted as unambiguous and definitively correct in a short format with minimal scoring variance; 4) Temporal grounded required: Requiring answers to rely on one or more temporal segments in videos, rather than single frames. We extensively evaluate 33 state-of-the-art LVLMs and summarize key findings as follows: 1) Current LVLMs exhibit notable deficiencies in factual adherence, with the best-performing model o3 merely achieving an F-score of 66.3%; 2) Most LVLMs are overconfident in what they generate, with self-stated confidence exceeding actual accuracy; 3) Retrieval-augmented generation demonstrates consistent improvements at the cost of additional inference time overhead; 4) Multi-hop QA demonstrates substantially degraded performance compared to single-hop sub-QAs, with first-hop object or event recognition emerging as the primary bottleneck. We position Video SimpleQA as the cornerstone benchmark for video factuality assessment, aiming to steer LVLM development toward verifiable grounding in real-world contexts.

VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering

Computation and Language

Helps computers answer questions about pictures more truthfully.

9 Mar 2025 2

90%

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency

CV and Pattern Recognition

Lets computers understand long videos by summarizing them.

1 Oct 2025 0

90%

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge

Computation and Language

Tests if AI tells the truth better.

9 Sep 2025 2

View PDF Login to Bookmark

Page Count

29 pages

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Tests if AI understands video facts correctly.

Technical Abstract

VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge