Score: 1

CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation

Published: August 10, 2025 | arXiv ID: 2508.07295v1

By: Yexing Du , Kaiyuan Liu , Youcheng Pan and more

Potential Business Impact:

Helps computers answer questions in many languages.

As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual modalities with a primary emphasis on English, which creates a gap in evaluation when processing multilingual input, especially in speech. To bridge this gap, we propose a novel \textbf{C}ross-lingual and \textbf{C}ross-modal \textbf{F}actuality benchmark (\textbf{CCFQA}). Specifically, the CCFQA benchmark contains parallel speech-text factual questions across 8 languages, designed to systematically evaluate MLLMs' cross-lingual and cross-modal factuality capabilities. Our experimental results demonstrate that current MLLMs still face substantial challenges on the CCFQA benchmark. Furthermore, we propose a few-shot transfer learning strategy that effectively transfers the Question Answering (QA) capabilities of LLMs in English to multilingual Spoken Question Answering (SQA) tasks, achieving competitive performance with GPT-4o-mini-Audio using just 5-shot training. We release CCFQA as a foundational research resource to promote the development of MLLMs with more robust and reliable speech understanding capabilities. Our code and dataset are available at https://github.com/yxduir/ccfqa.

CodeSimpleQA: Scaling Factuality in Code Large Language Models

Computation and Language

Tests if AI truly knows how to code.

22 Dec 2025 1

91%

SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

Computation and Language

Tests if AI answers questions about pictures correctly.

18 Feb 2025 1

91%

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks

Computation and Language

Tests AI that understands talking, seeing, and reading.

25 Jul 2025 2

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Repos / Data Links

github.com

Page Count

12 pages

CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation

Helps computers answer questions in many languages.

Technical Abstract

CodeSimpleQA: Scaling Factuality in Code Large Language Models

SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks