Score: 0

SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data

Published: November 17, 2025 | arXiv ID: 2511.17590v1

By: Ke Yu , Shigeru Ishikura , Yukari Usukura and more

Potential Business Impact:

Checks if fake data teaches computers right lessons.

Business Areas:

Semantic Search Internet Services

Synthetic tabular data, which are widely used in domains such as healthcare, enterprise operations, and customer analytics, are increasingly evaluated to ensure that they preserve both privacy and utility. While existing evaluation practices typically focus on distributional similarity (e.g., the Kullback-Leibler divergence) or predictive performance (e.g., Train-on-Synthetic-Test-on-Real (TSTR) accuracy), these approaches fail to assess semantic fidelity, that is, whether models trained on synthetic data follow reasoning patterns consistent with those trained on real data. To address this gap, we introduce the SHapley Additive exPlanations (SHAP) Distance, a novel explainability-aware metric that is defined as the cosine distance between the global SHAP attribution vectors derived from classifiers trained on real versus synthetic datasets. By analyzing datasets that span clinical health records with physiological features, enterprise invoice transactions with heterogeneous scales, and telecom churn logs with mixed categorical-numerical attributes, we demonstrate that the SHAP Distance reliably identifies semantic discrepancies that are overlooked by standard statistical and predictive measures. In particular, our results show that the SHAP Distance captures feature importance shifts and underrepresented tail effects that the Kullback-Leibler divergence and Train-on-Synthetic-Test-on-Real accuracy fail to detect. This study positions the SHAP Distance as a practical and discriminative tool for auditing the semantic fidelity of synthetic tabular data, and offers practical guidelines for integrating attribution-based evaluation into future benchmarking pipelines.

What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models

Machine Learning (CS)

Finds hidden problems in fake data.

29 Apr 2025 1

89%

SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot

Machine Learning (CS)

Shows why computers make certain decisions.

9 Oct 2025 0

87%

Causal SHAP: Feature Attribution with Dependency Awareness through Causal Discovery

Machine Learning (CS)

Shows what *really* makes computer guesses happen.

31 Aug 2025 0

View PDF Login to Bookmark

Country of Origin

🇯🇵 Japan

Page Count

10 pages

SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data

Checks if fake data teaches computers right lessons.

Technical Abstract

What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models

SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot

Causal SHAP: Feature Attribution with Dependency Awareness through Causal Discovery