| Speech enhancement(SE)methods have a widely used in real word,such as telephones,hearing aids,military communication devices,etc.Due to the development of artificial intelligence technologies,the intelligent voice assistants have earned great popularity in daily life,and place a higher demand on the capabilities to noise generalization.The addition of the SE module in the front-end of the speech recognition applications can help identify key instructions to a large extent.Methods based on deep neural networks contain three essential elements: input features,neural network,and training objects.As a source of information extraction for neural networks,input features play a crucial role in the entire network training process.In SE tasks,the commonly used features are in the temporal,frequency,and perceptual domains.In the time-frequency domain,it is difficult for neural networks to extract information as the coupling of speech and noise.In this paper,we propose a model for SE based on feature representations in the fractional domain.The work of this paper consists of two sections.(1)A deep neural network based on fractional domain features is proposed for SE.First,we analyze the distribution properties between speech and noise signals in the fractional plane.Second,the optimal order and the matrix order of the fractional spectrum of noisy speech are proposed as input features for the SE task.Then,the relevant experimental results demonstrate the effectiveness of the proposed model.(2)A dual path and interactive UNET for fractional domain-based SE is proposed.In this study,a fractional domain-based multi-order feature expression method is proposed by exploiting the difference between the aggregation of speech and noise in the fractional domain.The model is used to achieve synchronous estimation of speech and noise.The superiority of the proposed model is demonstrated by ablation experiments and performance experiments. |