K-means Algorithm Optimization and Parallel Computing Strategy Based on Flink Framework
LI Zhaoxin
MENG Xiangyin
XIAO Shide
HU Kaifeng
LAI Huanjie
Abstract:K-means algorithm is widely used in the field of machine learning and data mining because of its simple principle and good clustering effect,but it still has some shortcomings:K-means algorithm needs to specify the number of classification cate-gories K.K-means algorithm selection strategy for the initial clustering center is random selection,which may affect the accuracy and calculation speed of the final clustering results.The above shortcomings all limit the improvement of the calculation efficiency of the K-means algorithm.To solve the above problems,this paper proposes a K-means optimization algorithm based on Flink parallel-ization.This algorithm introduces the Canopy algorithm on the basis of the traditional K-means algorithm to complete the initial clus-tering,and obtains the number of categories K,and then uses the maximum distance algorithm to calculate the initial clustering cen-ter,and uses the parallel computing power of the Flink framework to perform clustering experiments on multiple data sets.The ex-perimental results show that the algorithm in this paper can reduce the number of iterations of the clustering process,and also has a certain improvement in the accuracy of clustering.It also has good computational efficiency in the environment of large-scale data sets.
Keywords:FlinkK-means algorithmCanopy algorithmparallel
Publication Date:2023-10-20
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:5( 2231-2235 )
