Research On Lightweight Speech Separation Method Based On Attention Mechanism | | Posted on:2024-07-06 | Degree:Master | Type:Thesis | | Country:China | Candidate:Z Y An | Full Text:PDF | | GTID:2568307157481364 | Subject:Master of Electronic Information (Professional Degree) | | Abstract/Summary: | | | Speech is one of the most important ways of human communication,which is often affected by noise and reverberation in indoor scenes.The goal of speech separation is to separate the target speech signal from these harsh environments.Traditional single-channel speech separation technology is usually based on the assumption that noise and speech are irrelevant.However,in an indoor reverberation environment,the target speech is affected by echoes.Microphone arrays can be used to collect mixed speech signals,and the phase information of speech signals can be obtained by using the inter-channel phase difference(IPD)between different microphone channels to assist speech signal separation.However,because the number of IPD features increases linearly with the square of the number of microphones,it is usually difficult to make full and effective use of all IPD feature information,which wastes some hardware resources in the microphone array.At the same time,with the development of deep learning,temporal convolutional networks(TCN)can save parameters,show better time series modeling ability than short-term memory networks,and show excellent performance in speech separation,but their parameters are still large.In view of the above problems,this paper takes the speech separation model based on the TCN network as the control group and makes the following research work:(1)In order to give full play to the channel resources of the microphone array and reduce the waste of hardware resources,this paper proposes an effective information acquisition method for IPD based on attention mechanisms.This method uses an attention scoring mechanism to generate the IPD weight of the next channel by dot product attention,so that IPD can be expressed in a reduced dimension without losing the phase information of the speech signal.This method can reduce the input load of the system when using IPD information with the same number of channels and can use as much IPD information as possible without increasing the input feature dimension,thus avoiding the waste of hardware resources on the microphone array.The TIMIT data set is used for experimental simulation.The experimental results show that the speech separation model of IPD effective information acquisition based on attention mechanism improves the scale-invariant signal-to-noise ratio by 3 d B,the model parameters are reduced by 36.3%,and the total floating-point number calculation of the model is reduced by 30.2% compared with the control group.(2)In order to further reduce the network parameters of TCN without degrading the speech separation performance.In this paper,a speech separation model based on lightweight TCN is proposed.According to the sparsity of speech signals in the timefrequency domain,the conventional extended convolution is selected to replace the deep extended convolution,and the gated branch is introduced to ensure the effective delivery of the target speech and greatly reduce the network parameters.The TIMIT data set is used for experimental simulation.The experimental results show that the speech separation model based on lightweight TCN improves the scale-invariant signal-to-noise ratio by 6.97 d B,the model parameters are reduced by 25.9%,and the total floating-point number calculation of the model is reduced by 24.7% compared with the control group. | | Keywords/Search Tags: | Reverberation environment, Speech separation, Microphone array, Attention mechanism, Gated branch, Lightweight | | Related items |
| |
|