Distributed Training Communication Optimization Method Based on Adaptive Hierarchical Gradient Compression
WANG Xiaoxiao
ZHU Xiaojuan
Abstract:To address the issues in the context of distributed machine learning of high communication overhead and low model training efficiency caused by frequent transmission of parameters and gradients between multiple computing nodes and parameter server nodes,a communication optimization method based on adaptive layered gradient compression(ALGC)was proposed.Firstly,an appropriate compression threshold was set for each layer of the neural network,and layers exceeding this threshold were selectively compressed.Secondly,a sparse threshold was separately set for each layer selected for compression and dynamically adjusted to achieve adaptive compression of gradient transmission for each layer.Finally,computation and communication were overlapped,and the parameter server aggregates the gradients and gradient residuals of each layer to update the global model.The results showed that the training accuracy of the ALGC method could reach up to 95.07%,and it achieved the minimum convergence time and the maximum speedup ratio.The ALGC method played a significant role in improving the model training speed and reducing communication overhead while ensuring the model training accuracy.
Keywords:distributed machine learninggradient compressionparameter serversparsificationcommunication optimization
Publication Date:2025-03-19
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:7( 34-40 )
