Research on mining short text information in E-government based on incremental clustering
LENG Yonglin
GUO Ying
SUN Xiaohong
QU Peiyi
Abstract:The E-government platforms generate a large amount of short texts every day.Clustering short texts play a very important role for the government's control of public opinion.This paper proposes a hybrid vector representation model that integrates weights and topic features to address the problem of feature information loss caused by a single short text vector representation model with limited information content.This model utilizes Word2vec and TF-IDF algorithms to mine local features of short texts,and utilizes BTM topic model to mine global features of short texts.Then,the two feature vectors are connected to form a short text vector.For the incremental changes of short texts,the Single-Pass clustering algorithm is improved by adding limited thresholds to achieve incremental clustering of short texts.The experimental results show that the hybrid vector representation model proposed in this paper can effectively improve clustering performance.
Keywords:E-governmentshort textvector representation modelincremental clustering
Publication Date:2023-09-15
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:8( 262-269 )