Score: 1

Enhancing Domain Diversity in Synthetic Data Face Recognition with Dataset Fusion

Published: July 22, 2025 | arXiv ID: 2507.16790v1

By: Anjith George, Sebastien Marcel

Potential Business Impact:

Creates better fake faces for computer training.

Business Areas:

Facial Recognition Data and Analytics, Software

While the accuracy of face recognition systems has improved significantly in recent years, the datasets used to train these models are often collected through web crawling without the explicit consent of users, raising ethical and privacy concerns. To address this, many recent approaches have explored the use of synthetic data for training face recognition models. However, these models typically underperform compared to those trained on real-world data. A common limitation is that a single generator model is often used to create the entire synthetic dataset, leading to model-specific artifacts that may cause overfitting to the generator's inherent biases and artifacts. In this work, we propose a solution by combining two state-of-the-art synthetic face datasets generated using architecturally distinct backbones. This fusion reduces model-specific artifacts, enhances diversity in pose, lighting, and demographics, and implicitly regularizes the face recognition model by emphasizing identity-relevant features. We evaluate the performance of models trained on this combined dataset using standard face recognition benchmarks and demonstrate that our approach achieves superior performance across many of these benchmarks.

Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data

CV and Pattern Recognition

Makes face recognition fairer with fake pictures.

28 Jul 2025 1

90%

Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise

CV and Pattern Recognition

Creates fake faces for computer recognition.

20 Oct 2025 0

90%

Model Discrepancy Learning: Synthetic Faces Detection Based on Multi-Reconstruction

CV and Pattern Recognition

Finds fake faces made by computers.

10 Apr 2025 1

View PDF Login to Bookmark

Page Count

7 pages

Enhancing Domain Diversity in Synthetic Data Face Recognition with Dataset Fusion

Creates better fake faces for computer training.

Technical Abstract

Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data

Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise

Model Discrepancy Learning: Synthetic Faces Detection Based on Multi-Reconstruction