Score: 0

DOD: Detection of outliers in high dimensional data with distance of distances

Published: November 4, 2025 | arXiv ID: 2511.02199v1

By: Seong-ho Lee, Yongho Jeon

Potential Business Impact:

Finds strange data points in complex information.

Business Areas:

Big Data Data and Analytics

Reliable outlier detection in high-dimensional data is crucial in modern science, yet it remains a challenging task. Traditional methods often break down in these settings due to their reliance on asymptotic behaviors with respect to sample size under fixed dimension. Furthermore, many modern alternatives introduce sophisticated statistical treatments and computational complexities. To overcome these issues, our approach leverages intuitive geometric properties of high-dimensional space, effectively turning the curse of dimensionality into an advantage. We propose two new outlyingness statistics based on observation's relational patterns with all other points, measured via pairwise distances or inner products. We establish a theoretical foundation for our statistics demonstrating that as the dimension grows, our statistics create a non-vanishing margin that asymptotically separates outliers from non-outliers. Based on this foundation, we develop practical outlier detection procedures, including a simple clustering-based algorithm and a distribution-free test using random rotations. Through simulation experiments and real data applications, we demonstrate that our proposed methods achieve a superior balance between detection power and false positive control, outperforming existing methods and establishing their practical utility in high-dimensional settings.

A method for outlier detection based on cluster analysis and visual expert criteria

Machine Learning (CS)

Finds weird data points hidden in big groups.

27 Oct 2025 0

87%

Dissecting Mahalanobis: How Feature Geometry and Normalization Shape OOD Detection

Machine Learning (CS)

Makes AI better at spotting fake or unusual things.

17 Oct 2025 1

87%

Contributions to Robust and Efficient Methods for Analysis of High Dimensional Data

Statistics Theory

Finds important patterns in huge, messy data.

9 Sep 2025 1

View PDF Login to Bookmark

Country of Origin

🇰🇷 Korea, Republic of

Page Count

29 pages

DOD: Detection of outliers in high dimensional data with distance of distances

Finds strange data points in complex information.

Technical Abstract

A method for outlier detection based on cluster analysis and visual expert criteria

Dissecting Mahalanobis: How Feature Geometry and Normalization Shape OOD Detection

Contributions to Robust and Efficient Methods for Analysis of High Dimensional Data