Research of Data Cleaning Algorithm in Data Warehouse
CAI ZHong-jie
LEI Bin
ZHANG Wei
Abstract:This paper describes some advices for improving the problems in the algorithm of duplicate elimination. The improved duplicate elimination algorithm has effectively promoted the efficiency of scheduling records on the environment that record matching rate was keeping high. In detecting duplicate records, it takes into account 5 factors. For instance, the number of characters ,the frequency of character be found in the 2 ifelds, the importance (weight)of ifeld in records, the Chinese semantic and these mantic focus is always in the back location etc;In merging duplicate records, it uses both the cluster algorithm and practical algorithm to do that. It makes the data cleaning algorithm more accurate and healthier.
Keywords:Scheduling recordDuplicate eliminationDetecting duplicate records
Publication Date:2013-01-01
Online Publishing Date:2026-05-22(First online date of this platform, not the publication date of the document)
Pages:4( 32-34,40 )
