Font Size: a A A

Research On Key Issues Of Speech Enhancement Algorithm Using Neural Network

Posted on:2022-10-04Degree:MasterType:Thesis
Country:ChinaCandidate:J C YuFull Text:PDF
GTID:2518306338467934Subject:Electronics and Communications Engineering
Abstract/Summary:
Speech enhancement task can be divided into two categories:interference suppression and source separation.It is one of the key research directions in the field of speech signal processing,and also one of the key front-end technologies of natural language processing,which has important research value.Because the assumption of traditional speech enhancement algorithm on signal limits its application scenario,neural network algorithm with strong generalization ability has become the mainstream algorithm.Therefore,this thesis mainly focuses on the neural network-based speech enhancement algorithm to carry out a series of research.1)In order to solve the problem of interference suppression,the thesis takes time domain convolutional neural network as the infrastructure,and focuses on the influence of masking mechanism,optimization criteria,residual block structure on the performance of the standard interference suppression neural network.By combining one-dimensional inversion bottleneck layer and convolution algorithm with holes,the basic interference suppression model is given.At the same time,by analyzing the optimization direction provided by the current mainstream loss function,the optimization criterion of angle distance based on waveform is proposed.The comparison results show that compared with the baseline model,the proposed interference suppression model has a scale invariant SNR gain of 0.2 db.2)In order to improve the reconstruction speech quality of the proposed interference suppression model,an adaptive neural network is formed by combining the reference information with the model.The network consists of two parts:skeleton network and reference information extractor.The interference suppression system composed of the two has the adaptive characteristics,and can adjust some features automatically according to the characteristics of input speech,and the two are used for interference suppression.The experiment shows that the adaptive network has a scale invariant SNR gain of 0.38 dB compared with the basic interference suppression model with the same parameter scale.In PESQ and STOI,the proposed model has 0.08 and 0.87 gains compared with the baseline model.3)In view of the unknown number of speakers and variable task of sound source separation,this thesis studies the algorithm of sound source counting and separation based on adaptive neural network,and realizes the sound source separation system with variable output.The proposed model is a reference information extractor which can generate the embedded feature and the center of mass of sound source at the same time-frequency level to avoid the clustering operation in the deep clustering algorithm.Meanwhile,the momentum contrast training algorithm is proposed to make the invalid centroid gather and remove according to the threshold value by using the differentiable characteristics of the centroid estimation module.The simulation results show that the proposed system can solve the problem of sound source counting to some extent,and the accuracy of the number estimation of the sound sources reaches 95.71%,and the output voice quality has a 1.12 dB scale invariant SNR improvement compared with the baseline model.
Keywords/Search Tags:Speech enhancement, Deep learning, Adaptive neural network, Source counting
Related items