Score: 0

Case Count Metric for Comparative Analysis of Entity Resolution Results

Published: January 6, 2026 | arXiv ID: 2601.02824v1

By: John R. Talburt , Muzakkiruddin Ahmed Mohammed , Mert Can Cakmak and more

Potential Business Impact:

Compares how two computer sorting methods work.

Business Areas:
CMS Information Technology, Software

This paper describes a new process and software system, the Case Count Metric System (CCMS), for systematically comparing and analyzing the outcomes of two different ER clustering processes acting on the same dataset when the true linking (labeling) is not known. The CCMS produces a set of counts that describe how the clusters produced by the first process are transformed by the second process based on four possible transformation scenarios. The transformations are that a cluster formed in the first process either remains unchanged, merges into a larger cluster, is partitioned into smaller clusters, or otherwise overlaps with multiple clusters formed in the second process. The CCMS produces a count for each of these cases, accounting for every cluster formed in the first process. In addition, when run in analysis mode, the CCMS program can assist the user in evaluating these changes by displaying the details for all changes or only for certain types of changes. The paper includes a detailed description of the CCMS process and program and examples of how the CCMS has been applied in university and industry research.

Country of Origin
🇺🇸 United States

Page Count
17 pages

Category
Computer Science:
Databases