Text data mining algorithm for multi-label implicit knowledge
DENG Qiaofu
LI Xiaoya
GUO Xiaojun
Abstract:[Objective]With the expanding user group of social software,multi-label annotation has been increasingly adopted for text information.How to analyze the behavior and psychology of the user group through data mining of multi-label text information has become a research hotspot.A data mining algorithm for multi-label implicit knowledge based on a deep topic feature extraction model was utilized to enhance text classification accuracy and data mining efficiency.[Methods]To deeply understand the implicit knowledge in text information,the socialization,externalization,combination,and internalization(SECI)theory was employed to convert the implicit knowledge into explicit knowledge.The short-term memory capability of recurrent neural networks was utilized to improve the conversion efficiency.Considering the complexity of text information,local and global features were analyzed separately,and feature fusion was used to improve data mining efficiency.Due to the strong correlation between the context of text information,the gate mechanism of the long short-term memory(LSTM)model was applied to extract contextual dependencies,while the unsupervised latent Dirichlet allocation(LDA)topic model was selected to model the topic structure of the text to mitigate standard differences from manual labeling.Combining LDA-derived global features and LSTM-derived local features,feature stitching was performed to reduce information loss during the feature extraction.A theme controller was introduced to narrow down the inference scope,which obtained more effective text features.Simultaneously,a Gaussian decoder-based contextual topic layer was constructed to calculate the conditional probability matrix of each vocabulary under a given topic,and a Gaussian mixture decoder was used to obtain the conditional probability of the vocabulary.Topic modeling optimization and content expansion were achieved through a Gaussian mixture decoder.Finally,multi-label classification was implemented using the Softmax function to calculate label probabilities.[Results]During model training,perplexity was used as a criterion for evaluation.The proposed model exhibited better perplexity than the control groups(LDA topic model and LSTM model),demonstrating the effectiveness of feature concatenation combining the LDA topic model and LSTM model.By comparing with NVDM,LSTM,LDA,and VAETM models,with precision and recall as evaluation metrics,the proposed model improves precision and recall by 5.05%and 2.75%,respectively.[Conclusion]The comparative experimental results show that the proposed model can significantly improve the performance of text classification.Compared with the LDA topic model and the LSTM model,it outperforms in processing multi-label texts.It can efficiently mine the implicit knowledge in multi-label text data,providing an efficient and accurate solution for tasks such as text classification,semantic analysis,and information retrieval.
Keywords:multi-label textdeep topic feature extraction modelimplicit knowledgerecurrent neural networklong short-term memory(LSTM)neural networklatent Dirichlet allocation(LDA)topic modelfeature stitchingGaussian decoder
Publication Date:2025-09-25
Online Publishing Date:2025-10-31(First online date of this platform, not the publication date of the document)
Pages:8( 594-601 )
