Detecting Approximate Duplicate Records Based on K-modes Clustering Algorithm
CHEN Yanping
HONG Mingjie
YANG Xiaobao
Abstract:Detecting approximately duplicated records has become an important branch of data cleaning and it's an important way in eliminating data redundancy to improve the data quality,used in the data statistics,data analysis,data warehouse,artificial intelligence and data mining. This paper studies the current approximately duplicated records detection method. For there are many methods to detect the problem of low detection accuracy and efficiency,using K-Modes to cluster method and information entropy theory to reduction of dimensions and determine the attribute weights. In the record matching phase,records are compared accord?ing to the importance of the attribute,and the approximately duplicated records are judged according to the threshold value. This method avoids the whole record of matching and saves time. After each data set is completed,the approximately duplicate records are finally eliminated. The experiment shows that this method can effectively reduce the detection data set range and detection effi?ciency,improve the time efficiency and detection accuracy,and have higher detection rate and precision.
Keywords:approximate duplicate recordK-Modes clustering algorithminformation entropysimilar detecting
Publication Date:2019-01-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:7( 2966-2972 )
