Font Size: a A A

Research On Protein Hot Spot Residues Prediction Based On Encoding Of Sequence-Segment Neighbors

Posted on:2020-07-27Degree:MasterType:Thesis
Country:ChinaCandidate:T ShenFull Text:PDF
GTID:2370330575971068Subject:Biology
Abstract/Summary:
When a protein interacts with a protein,its free energy of binding is contributed by only a small portion of the amino acid residues,which are called hot spot residues.The protein function is frequently dependent on hot spot residues.Hot spot residues usually tend to accumulate at the center of the protein interaction interface.They play vital roles in protein binding.Therefore,we need to deepen the understanding for hots pot residues.It is significant on the development of life sciences.At present,researchers rely mainly on Alanine scanning mutagenesis technology to identify hot spot residues,but this method is costly,time-consuming and laborious.It can only be applied in a small range.Therefore,we urgently need a more accurate and efficient method to identify protein interface hot spot residues.This paper puts forward encoding of sequence-segment neighbors method and constructs a prediction model based on Random Forest classification algorithm to identify hot spot residues in the protein interaction interface.First,The training set was extracted from the ASEdb database.Then 10 amino acid physicochemical properties,16 features related to the PI and DI,and 25 features related to ASA were extracted.The most important thing is that we improve the protein coding method and provides a new idea for the prediction protein hot spot residues.Different from the previous methods of using auto correlation descriptor or triplet combination information to code protein sequences,it is realized in this paper that the influence of amino acids adjacent to hot spot residues and amino acids that have a certain interval with hot spot residues.We adjust the length of the sliding window.And protein sequence is divided into 3,4,and 5 segments.The prediction model is established.The optimal setting parameters are finally selected through cross-validation.To verify the reliability of the prediction model,this paper extracted the independent test set from the BID database to validate the model.Finally,the prediction model is compared with the existing hot spot residues prediction methods.These models are of great significance in this research direction,including APIS,Robetta,FOLDEF,KFC and MINERVA models.Among the models constructed by the same training set,the model of this paper significantly improved the prediction ability of protein interface hot spot residues on the same test set,It indicate that the method is reliable.
Keywords/Search Tags:Protein interaction, Hot spots, Encoding of sequence-segment neighbors, Sliding window, Random Forest
Related items