Multi-source Feature Decoupling Voice Conversion Model Based on Diffusion Models
ZHANG Yedong
WEN Shuangbing
TAN Hanzhong
HUANG Haifeng
HU Tao
Abstract:To address issues such as timbre leakage,loss of prosodic details,and insufficient naturalness in generated speech,a multi-source feature decoupling voice conversion model based on diffusion models(MFD-VC)was proposed.Firstly,speech was decomposed into distinct subspaces with different attributes and was processed independently in the model,through which high-fidelity voice conversion was achieved via multi-attribute collaborative control.Secondly,a style encoder(SE)was designed to extract speaker timbre features from reference speech.Concurrently,a content feature encoder(CFE)and a wave network-based encoder(WN)were utilized to process content and fundamental frequency information,respectively.Finally,a multi-scale fusion 1 dimension(MSF1D)module was integrated into the U network diffusion model(UNetDiff)architecture to enhance multi-scale feature expression in skip connections,thereby allowing the generation of more detailed Mel-spectrum representations.The results showed that the MFD-VC achieved an equal error rate of only 12.2%on the Librispeech for text-to-speech(LibriTTS)dataset,while a high average opinion score of 4.35 for similarity was obtained.Moreover,on the voice cloning toolkit(VCTK)dataset,an equal error rate of only 15.2%and a high average opinion score of 4.24 for similarity were achieved.The MFD-VC delivered superior performance in terms of sound quality,similarity,and content clarity.
Keywords:voice conversionMFD-VCdiffusion modelwave networkUNetDiff
Publication Date:2026-03-20
Online Publishing Date:2026-03-26(First online date of this platform, not the publication date of the document)
Pages:5( 57-61 )