Font Size: a A A

Research On The Method Of Speaker-specific Speech Signal Recognition

Posted on:2024-05-23Degree:MasterType:Thesis
Country:ChinaCandidate:J X LiFull Text:PDF
GTID:2568307079965849Subject:Electronic information
Abstract/Summary:
Speaker specific speech signal recognition(SSSR),also known as speaker recognition(SR),as the name implies,is the identification of speakers through sound features,which belongs to the development of biometric technology in the field of sound.Speaker recognition is a branch of many biometric technologies,which can be divided into two categories: speaker recognition and speaker verification.Among them,speaker recognition is a one to multiple task: it will determine the current speaker identity from speakers registered in the speech database;Speaker confirmation belongs to a one-on-one problem: give a person’s voice characteristics,and determine whether they are the same person based on their previous registered voice.The technical basis for the realization of speaker recognition tasks is the uniqueness of voice,which means that different people have different speech characteristics,such as fingerprints.With this unique feature,the task of recognizing the identity of different people can be realized based on voice information.The main work of this article is as follows:1.For TDNN networks,this thesis introduces a residual structure to form an RTDNN network,enabling it to learn better and richer speech features.2.For R-TDNN networks,this thesis introduces an attention mechanism,conducts simulation experiments on R-TDNN networks based on two attention mechanisms,and compares their final model performance.3.Based on the introduction of attention mechanism,this thesis improves the ECAPA-TDNN network,and conducts experimental simulations on the improved ECAPA-TDNN network.The experiments show that the improved network performs better than the original network.4.For input data,this article has conducted data enhancement of voice data,and conducted simulation experiments to prove that data enhancement indeed helps optimize the model.5.For the speaker recognition models obtained after training,they were divided into two scenarios: noisy and non noisy.Speaker recognition tasks were conducted for these two scenarios,and the number of times each model could correctly recognize was counted out of 100 recognition times.
Keywords/Search Tags:Speaker Recognition, Attention Mechanism, Neural Network, Data Enhancement
Related items