Data cleaning technology based on N-Gram algorithm
MA Ping-quan
SONG Kai
JI Jian-wei
Abstract:Aiming at the plentiful approximately duplicate data in the database,the attribute structure of approximately duplicate records and the causing reason were analyzed.The data records were calculated with the N-Gram algorithm to get the key values,namely N-Gram values,which represented the attribute of every record.According to the key values,the data records in the database were ordered so as to form a well-organized database.In addition,the similarity of data records in the database was calculated.The identified approximately duplicate records were cleaned by applying the arranged combination cleaning idea. The experimental results show that the N-Gram algorithm effectively increases the recall ratio and precision ratio of approximately duplicate data records.
Keywords:similarityapproximately duplicate recordattributeorderingcombinationdata cleaningrecall ratioprecision ratio
Publication Date:2017-01-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:6( 67-72 )
