Font Size: a A A

Study On Speech Separation In A Multi-signal Source Scenarione

Posted on:2024-02-12Degree:MasterType:Thesis
Country:ChinaCandidate:Q L JiangFull Text:PDF
GTID:2568307184956349Subject:Master of Electronic Information (Professional Degree)
Abstract/Summary:
In recent years,due to the development and progress of deep learning and natural language processing,speech separation technology has also ushered in new opportunities and challenges along with the development of artificial intelligence.The widespread use of voice in daily life makes the intelligent voice technology also involved in all aspects of life applications,such as voice print wake-up,remote meeting,human-computer voice interaction and so on.But in daily life,the obtained sound signals obtained are not all single,more mixed with the interference of other sound signals.In this case,how to more effectively separate the required speech information from the mixed speech signals is a question worth studying.In addition,most of the speech separation models sacrifice a lot of computing power and resource costs along with the hierarchical superposition of the network.At the same time,considering that for the application of speech separation technology to daily life,the model needs to be applied to portable mobile devices with low power consumption,the lightweight of speech separation model is also a research problem to be implemented.To this end,this thesis studies the speech separation technology for multiple signal source scenarios,as follows:(1)In view of the complexity of model computation caused by the expansion network of speech separation model,Unet structure is introduced into the time domain convolutional network model,so that the model can obtain a small amount of computation.Meanwhile,in order to extract the loss of reduced information without improving the network calculation amount,this thesis,ECA attention mechanism and downsampling are adopted in the Unet structure.(2)For how to realize accurate speech signal separation from multiple signal source mixed voice,this thesis introduces the double path Transformer network structure voice print embedded coding mechanism,through the single speaker voice signal input to the encoder,get its voice feature coding,then voice coding embedded in the separator,can separate the speaker voice from the mix.By repeating the above operations and performing successive iterative separation,the speech separation problem for multiple signal sources can be solved.At the same time,considering that the dual-way Transformer network needs to learn a large number of parameters,so the computational cost of the model is very large,this thesis uses the resource efficient separation converter to realize the lightweight optimization of the model.Experiments were compared on the open-source dataset Mozilla Common Voice dataset(MCVD).The experimental results show that the improved separation model based on the time-domain convolutional network can separate the mixed speech of two speakers,and the number of parameters is reduced by 89.7%,and the separation result SI-SDR is14.1d B,which is 2.4d B higher than the baseline model,indicating that this improved model can achieve good separation effect with a very small number of parameters.The speech separation model based on Transformer network can realize the mixed speech separation of three or more speakers,and optimize the lightweight module of the model to obtain the separation effect of SI-SDR=15.5d B when the number of parameters is reduced by 70%,which improves by 1.1d B compared with the dual-path Transformer network.At the same time,it explores the difference of mixed speech signal separation effect under different language conditions,and draws the conclusion that different language conditions will affect the separation performance of the model.
Keywords/Search Tags:Speech separation, Self-attention mechanism, Dual-path network, Voice print embedded in the encoding mechanism, Lightweight
Related items