| Research background and purpose:Breast cancer has been recognized as the most common malignant tumor in female cancer patients worldwide,and it is usually the main cause of cancer-related deaths in women in developed and developing countries.The World Health Organization(WHO)annual evaluation report on cancer shows that in 2012,there were an estimated 1.7 million breast cancer patients worldwide,accounting for 25%of all cancer cases,of which 521,900 breast cancer-related deaths,accounting for all 15%of cancer deaths.In many countries in South America,Africa and Asia,the average annual incidence of breast cancer has been increasing,and it is gradually becoming younger[1].According to the analysis of research data in China,among women,breast cancer is the most common cancer between the ages of 30 and 59,and is the main cause of cancer death in women under 45[2].With the further increase in the number of new cases each year,breast cancer has gradually developed into the most important health care problem and economic burden in my country and even in the world(Montero et al.,2012;Tao et al.,2015).Therefore,given that cancer prevention and control rely on population-based morbidity and mortality data,we should take action and evaluate current interventions to develop more effective breast cancer diagnosis and treatment strategies.Breast cancer is a highly specific tumor,and its treatment and prognosis are related to many factors.So far,the known factors that affect the early treatment,clinical surgery and prognosis of breast cancer mainly include the patient’s age,tumor size,lymph node metastasis,and histological grade.The expression levels of estrogen receptor(ER),progesterone receptor(PR),human epidermal growth factor receptor-2(HER-2),and Ki-67 protein at a certain cellular molecular structure level also play a irreplaceable role in the prognosis of breast cancer.And with the rapid development of advanced precision medicine,high-throughput sequencing integration technology and genome detection chip technology,more and more scholars have turned their attention to the field of breast cancer molecular therapy.Therefore,studying the molecular biology basis of the early occurrence and progression of breast cancer,discovering the corresponding diagnostic and therapeutic molecular markers,and identifying new breast cancer prognostic biomarkers will help predict its biological behavior and build a useful guide for clinical diagnosis and treatment.It is a vital tool for predicting the prognosis of breast cancer patients,and helps to improve the design of individualized treatment plans and develop new therapeutic targets.The expression of DNA damage and damage repair genes is related to the occurrence and biological behavior of various tumors,suggesting its potential as prognostic markers and therapeutic targets.There have been different reports on the prognostic value of DNA damage repair gene expression in breast cancer.In this study,we used and integrated the transcriptome information and clinical data of breast cancer in the TCGA database(The Cancer Genome Atlas(TCGA))to analyze the differentially expressed genes in breast cancer samples and normal samples to construct A clinical prognostic risk model closely related to breast cancer DNA damage repair genes,explore the expression and clinical therapeutic value of DNA damage repair genes in breast cancer,and verify the predictive value of this model in overall breast cancer patients,so as to find new Breast cancer targeted therapy provides certain reference value.Method:Downloads the Manifest and Metadata data of the TCGA-BRCA transcriptome through The Cancer Genome Atlas(TCGA)website,and then downloads the original HTSeq-Counts data in the cmd environment with the help of the GDC-client download tool,using The Perl language script extracts the expression matrix of the original data.Downloads the Homo_sapiens.GRCh38.95.chr.gtf.gz file from the Ensembl website,and obtains the gene expression profile matrix based on the gene symbol after comparison;Usees the"limma" package of the R language to screen the differentially expressed genes(DEGs)of breast cancer and normal breast mRNA expression data,and the screening conditions were set to(|logFC|>1.0 and the adjusted pvalue,FDR<0.05);then,on the one hand,use the David website(https://david.ncifcrf.gov/tools.jsp)and the KOBAS website(http://kobas.cbi.pku.edu.cn/)respectively perform GO function enrichment analysis to obtain differential breast cancer DNA damage repair genes,and use Cytoscape and R software to visualize the results.Combine the DNA damage repair gene collection obtained by the two methods of the David website and the KOBAS website.On the other hand,through the Amigo2 database(http://amigo.geneontology.org/amigo/landing)download the DNA damage repair gene set numbered GO.0006281,and use the R language"colorfulVennPlot" package to process the downloaded gene set and differential genes Obtain differential breast cancer DNA damage repair genes.Finally,the DNA damage repair genes obtained from the two aspects are integrated and further analyzed for KEGG pathway enrichment.At the same time,download the clinical survival data of TCGA-BRCA from the TCGA database,use the R language script to merge the survival data and the differential DNA damage repair-related gene expression data,and perform the univariate COX proportional hazard regression model analysis,and then based on the the P value of single factor selects DNA damage repair genes related to survival prognosis for subsequent multivariate COX regression analysis.Construct a survival-related linear risk assessment model based on the expression profile and regression coefficients of the selected DNA damage repair genes after multi-factor COX regression analysis,calculate the risk score of each sample,and take the median of the risk score as the cut-off value.It is used to divide the samples into high and low risk groups;The time-dependent ROC curve is used to evaluate the predictive ability of the prognostic model in the 5-year survival period,and the Kaplan-Meier method is further used to draw the survival curves of the high and low risk groups.Use R language random sentences to divide the overall sample into two parts:"test group" and "train group".The samples of test group and train group are independent of each other.Repeating the above statistical method to calculate the risk value of each sample in the two groups of samples(risk score),according to the median value of risk score,each subgroup is divided into high and low risk groups.Survival analysis and ROC curve are used to analyze each subgroup to further verify the reliability of the prognostic risk model.The results are visualized using R software.Results:A total of 1222 samples of transcriptome counts data were obtained from the TCGA database,including 113 normal samples and 1109 tumor samples.After integration,56753 gene expression profile matrices were obtained.At the same time,the downloaded clinical data were processed to obtain clinical data of 1085 female breast cancer patients.After screening for differential genes,a total of 4177 differentially expressed genes were obtained,of which 2247 were up-regulated and 1930 were down-regulated.The 112 differential breast cancer DNA damage genes obtained through the analysis of the David,KOBAS and Amigo2 websites were subjected to single-factor-COX regression analysis.After the P value was less than 0.05,a total of 18 differential genes related to prognosis were screened,including RAD54B,RAD21,PARPBP,BRCA1,TIMELESS,CLSPN,CHEK1,CHAF1B,FANCD2,BRCA2,RAD51,MCM4,EME2,HIST3H2A,GINS4,MCM6,CDCA5,PYCARD.Among them,15 differential genes(RAD54B,RAD21,PARPBP,BRCA1,TIMELESS,CLSPN,CHEK1,CHAF1B,FANCD2,BRCA2,RAD51,MCM4,GINS4,MCM6,CDCA5)were negatively correlated with patient survival time,and 3 genes(HIST3H2A,HIST3H2A,PYCARD,EME2)are positively correlated with patient survival time.The expression and clinical data matrix of 18 prognostic-related differential genes were reconstructed to perform multi-factor COX regression analysis,and 4 genes significantly related to prognosis were screened out:GINS4,RAD54B,BRCA1,EME2.The regression coefficients of the multi-factor COX analysis of the 4 differential genes were further extracted,and the risk value of each sample was calculated,and a prognostic risk scoring model composed of these 4 genes was constructed.The prognostic score(PI)formula is:PI=-0.14502×GINS4 expression+0.43840×RAD54B expression+0.16469×BRCA1 expression-0.24295×EME2 expression.After calculating the prognosis score of 1078 patients,the median value was 0.978.A total of 539 patients of 1078 patients were included in the high-risk group,and 539 patients were included in the low-risk group.Use R language to draw high and low risk heat maps,ROC curves and KM survival curves.The time-dependent ROC curve shows that the risk assessment model has certain significance in predicting the 5-year survival prognosis of breast cancer patients(the area under the ROC curve of 5-year survival rate AUC is 0.657).The K-M survival curves of samples from the high-risk and low-risk groups indicated that the overall survival rate of patients in the high-risk group was lower,and the difference between the two groups was statistically significant(P=0.00077).The K-M survival curves of the test group and the train group also showed that the overall survival rate of patients in the high-risk group was lower,and the difference between the two groups was statistically significant(P=0.04525,P=0.00416,respectively).The ROC curves of the two subgroups showed that the 5-year survival rate AUC was 0.654 and 0.605,respectively,indicating that the model has a certain degree of stability and validity.Conclusion:The risk prognosis model constructed based on breast cancer DNA damage repair genes can predict the survival and prognosis of breast cancer patients at a certain level,and has certain reference and reference value for the prognosis of breast cancer patients.Combining the prognostic factors at the molecular structure level of breast cancer cells can further screen out high-risk breast cancer groups and guide the formulation of more effective individualized treatment plans. |