Long text clustering model based on contrastive learning optimization
WANG Yicheng
ZUO Weibing
Abstract:Long text clustering faces the challenge of difficult semantic representation.In order to meet this challenge and improve the performance of text clustering,this paper proposes a long text clustering model based on contrastive learning optimization(CLSSK).This model constructs a contrastive learning framework,embeds the SBERT pre-trained model improved by secondary pooling as a long text encoder,introduces the NT-Xent loss function to optimize the parameters of the long text encoder model,uses the optimized encoder to represent the long text,and finally inputs the obtained long text feature vector into the K-Means++algorithm for clustering analysis.The long text clustering experimental results show that the proposed model is better than the contrast model and the ablation variants.Therefore,the proposed model and its design modules have proven to be effective.
Keywords:long-text clusteringCLSSK modelcontrastive learningSBERTK-Means++
Publication Date:2025-12-25
Online Publishing Date:2026-01-27(First online date of this platform, not the publication date of the document)
Pages:7( 24-30 )