Font Size: a A A

Proteogenomics Study On Small Proteome Of Saccharomyces Ceresiciae

Posted on:2019-05-07Degree:MasterType:Thesis
Country:ChinaCandidate:C T HeFull Text:PDF
GTID:2370330545963166Subject:Biochemistry and Molecular Biology
Abstract/Summary:
Small proteins(SPs)are defined as peptides of 100 amino acids(AA)or less encoded by short open reading frames(sORFs).It has been demonstrated that SPs participate a wide range of function in cells,including gene regulating,cell signaling,metabolism,etc.But the progress of SPs is slow because of many technical problems.The annotation of SPs is challenging because of their small molecular weight.Most of the annotated SPs in all living organisms are currently lacking expression evidence at the protein level,and regarded as missing proteins(MPs).The discovery of more and more unannotated sORFs also reveals a blind spot in traditional gene annotation for these genes.Many challenges have been found in SPs identification because of low abundant by proteomics.The deeper coverage of SPs identification will help to validate more MPs and discover more unannotated sORFs.Firstly,we optimized SPs enrichment strategies on-Saccharomyces cerevisiae S288 C.By integrating four different SPs enrichment methods,we identified 117 SPs,validated 31 MPs and successfully verified 3 novel sORFs(YKL104W-A,YHR052C-B and YHR054C-B)by novel peptide synthesis.Further analysis showed that the main missing factors of MPs were low molecular weight,low abundant,hydrophobicity,lower codon usage bias,unstable,etc.The small protein-based enrichment can be used for MPs and sORFs searching,which might provide the foundation for their function research.Secondly,we established proteogenomic flow for SPs.In our first part,all annotated and novel peptides were combined and filtered to obtain a more reliable identification set by a global false discovery rate(FDR)and three separated FDR(S-FDR,T-FDR I and TFDR II);we obtained and validated three novel peptides.Further analysis showed that three FDRs were too strict to ignore many confident novel peptides for small database.The sensitivity and accuracy cannot be easily considered simultantly based on FDR selection strategy.The high quality and low quality peptide-spectrum matching can’t be distinguished accurately by current peptide scores,which might the main factor of high false positive rate during novel peptide identification.To solve the problem,we developed a solution for filtering the low quality peptide-spectrum matching spectra.In addition,to ensure the sensitivity of the novel peptide identification,we applied an absolute threshold which correspond to the minimum standard of true positive novel peptide to all novel peptides for the preliminary screening.Finally,we established a novel peptide selection strategy which was dependent on the absolute threshold filtering based on Raw and PSM score.Totally,we identified and validated 7 novel sORFs(YBL071W-B,YDL191C-A,YHL029W-A,YMR106C-A,YBL063C-A,YKL038W-A and YPL049C-A)in S.cerevisiae S288 C.Homology analysis showed that these 7 novel genes were yeast unannotated genes,which also behaved lower similarity with other species.
Keywords/Search Tags:Proteogenomics, small protein, enrichment, missing proteins, sORF
Related items