Convolutional neural network compression method based on knowledge distillation
ZHENG Yun
GAO Peng
Abstract:[Objective]Convolutional neural networks,as an important technology in the field of deep learning,have demonstrated outstanding performance in multiple areas such as image recognition,object detection,and natural language processing.However,with the increase in model depth and complexity,the size and computational requirements of convolutional neural network models increase sharply,which poses severe challenges to model deployment and real-time applications.[Methods]Therefore,to reduce the size and computational complexity of neural network and improve the efficiency and deployability of the models,a convolutional neural network compression method based on knowledge distillation was proposed.Model compression and acceleration could be achieved by transferring knowledge from large complex models(teacher network models)to small simplified models(student network models).A high-performance teacher network and a student network with a simpler structure and fewer parameters were established.Abundant feature representations and accurate prediction results were provided by the teacher network,while the student network learned the behavior of the teacher network to approach its performance.A standard loss function was used,and its parameters were iteratively updated through the back propagation algorithm to ensure that it achieved good performance on the training dataset.An improved knowledge distillation method was adopted to obtain a comprehensive threshold function,which was used to evaluate the knowledge differences between the teacher network and the student network and guide the learning process of the student network.During the training process,the student network was supervised by using this comprehensive threshold function,gradually approaching the output of the teacher network while maintaining a small model size and computational complexity.In this way,the compression processing of the convolutional neural network was realized.[Results]The results show that the proposed method exhibits good model compression performance on both ImageNet and Labelme datasets.The proposed method has a high degree of fitting for the output results of the convolutional neural network before and after compression,which indicates that the student network has successfully learned the key features of the teacher network.The cross entropy loss value is relatively low,around 1.0,further verifying its good predictive performance.The compression time for completing the convolutional neural network model is relatively short,between 79.8 and 89.4 s,indicating that the proposed method has high computational efficiency.[Conclusion]From the above results,it can be seen that the convolutional neural network compression method based on knowledge distillation can effectively reduce model size,decrease computational complexity,and maintain or even improve model performance.The proposed method provides not only a new approach for model compression but also strong support for the deployment and application of deep learning models.In addition,the proposed method is improved on the basis of the knowledge distillation method.A comprehensive threshold function is introduced to more comprehensively evaluate and guide the learning process of the model,which enhances the effectiveness and efficiency of knowledge distillation to some extent.Therefore,the proposed method has not only theoretical value but also important practical significance.
Keywords:convolutional neural network compressionimproved knowledge distillation methoddiscriminatorstudent networkteacher networkstandard loss functioncomprehensive threshold functioncross entropy loss value
Publication Date:2025-05-25
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:7( 348-354 )
