Font Size: a A A

Research On Deep Learning Model For Speech Enhancement And Separation In Complex Auditory Scenes

Posted on:2023-06-24Degree:MasterType:Thesis
Country:ChinaCandidate:Y H WuFull Text:PDF
GTID:2558306914473754Subject:Electronic and communication engineering
Abstract/Summary:
Speech enhancement task aims to deal with those speech signals that are disturbed or even drowned by various noises,and extract useful speech signals by suppressing and reducing irrelevant noises,so as to restore pure original speech as much as possible.Speech enhancement as a key research area of speech signal processing has attracted extensive attention.Speech enhancement tasks in broad sense can be divided into noise suppression and sound source separation according to different application scenarios.In recent years,with the continuous development of neural networks,more and more neural networks with more generalization ability are used to enhance the performance of speech enhancement tasks.Speech enhancement algorithm based on neural network has become a key research topic in the field of deep learning.Therefore,this paper mainly focuses on speech enhancement algorithms based on deep neural networks.1)Aiming at the problem of noise suppression,the time-domain analysis framework is used as the basic research framework,and the convolutional neural network based on encoder-decoder structure is used for mapping modeling.The effect of feature processor design and processing block arrangement on speech enhancement performance of neural network is studied.Double-Stream module is proposed,in which two parallel asymmetric convolution blocks are set up and the feature interaction between the two streams is utilized to enhance the feature processing capability of the network.Experiments show that compared with the baseline model,the model proposed in this paper has achieved better results in most evaluation indicators,and has stronger learning ability and application potential.2)In order to solve the problem of insufficient ability of DoubleStream TCN network to integrate the characteristics of each tributary and further improve the speech reconstruction quality of the proposed speech model,a Double-Stream Gated Output Network(DSGON)is proposed.In the proposed network model,the Double-Stream Gated Output module recombines and splices the main stream and auxiliary stream features transmitted from the backbone network,and completes the integration of Double-Stream features through feature recombination and gated output module,and finally forms masking output.Experiments show that the Double-Stream Gated Output module has great integration ability to integrate the main stream and auxiliary stream characteristics of the Double-Stream network,and effectively improves the speech enhancement performance of the Double-Stream network.
Keywords/Search Tags:speech enhancement, adaptive neural network, deep learning
Related items