| Background:Lung cancer is one of the most commonly diagnosed malignancies in the world.According to the latest statistics from the International Agency for Research on Cancer(IARC),the incidence and mortality rate of lung cancer ranked the first among all cancers worldwide in 2018.In China,the incidence rate of lung cancer ranked the first among male malignancies and the second among female malignancies,and the mortality rate of lung cancer ranked the first among all cancers in both males and females.The high incidence and mortality rates of lung cancer in China can be attributed to the aging of Chinese populations and the high rate of tobacco use.Non-small cell lung cancer(NSCLC)is the major histological type of lung cancer,and accounts for 85%of incident lung cancer cases.The development of lung cancer involves the interplay between environmental and genetic factors.Smoking is the major risk factor for lung cancer.However,less than 20%of smokers will develop into lung cancer,which indicates that the susceptibility to lung cancer varies among different individuals.The varying susceptibility to lung cancer could be caused by germline variants.One of the common types of germline variants is single nucleotide polymorphism(SNP).With the rapid advancement of genomic technologies,genome-wide association study(GWAS)becomes an efficient research tool in molecular epidemiology,and has been widely used in exploring genetic risk loci for complex diseases or traits.Since 2008,multiple GWASs of lung cancer have been conducted in European and Asian populations.And a total of 45 susceptibility loci for lung cancer have been identified.Although GWASs have made great achievements in investigating the susceptibility loci for lung cancer,risk SNPs of lung cancer that were identified by GWASs only explained a small proportion of lung cancer heritability,which is called“missing heritability”.One of the explanations for the missing heritability is that much larger numbers of genetic variants have not yet been found.Recently,several consortiums were funded based on the cooperation of multiple research institutions around the world.These consortiums try to combine GWAS data from multiple research centers with the aim to identify additional susceptibility loci for complex diseases.With regard to research on cancer,The OncoArray Consortium was to gain new insight into the genetic architecture and mechanisms underlying breast,ovarian,prostate,colorectal,and lung cancers.In2017,The OncoArray Consortium published a lung cancer GWAS with the largest sample size to date in European populations,which identified 10 new susceptibility loci for lung cancer.In addition to conducting ancestry specific GWAS,GWASs are expanding to include multi-ethnic populations.In recent years,many novel disease-associated risk SNPs are being identified by genome-wide trans-ancestry meta-analysis.These results suggest that trans-ancestry meta-analysis can increase the sample size of GWASs,thus improving the statistical efficacy of the study.Aim:We aimed to identify new risk loci for non-small cell lung cancer(NSCLC).We also compared the effects of risk SNPs between Chinese and European populations.Gene-environment interaction analysis was conducted to identify risk SNPs that significantly interacted with smoking.We collected lung cancer risk SNPs of Chinese populations,and calculated the heritability that can be explained by these risk SNPs.And we also constructed polygenic risk score(PRS)to evaluate the utility of PRS in risk stratification for lung cancer.Methods:We conducted case-control studies that included a total of 27,120 NSCLC cases and 27,355 controls recruited from multiple research centers.All cases and controls were from seven GWASs of four studies:(1)the genome-wide association study of NSCLC in Chinese populations(NJMU GSA Project),which included the Nanjing GSA GWAS(4,149 cases and 3,198 controls),the Beijing GSA GWAS(2,155 cases and 2,035 controls),and the Guangzhou GSA GWAS(3,944 cases and4,065 controls);(2)the prior scanned lung cancer GWAS of Han Chinese(NJMU GWAS),which included the Nanjing GWAS(1,473 NSCLC cases and 1,962 controls)and the Beijing GWAS(858 NSCLC cases and 1,115 controls);(3)the lung cancer OncoArray GWAS of Chinese populations(the NJMU OncoArray GWAS with 953NSCLC cases and 953 controls);and(4)the lung cancer GWAS of European populations(TRICL-ILCCO OncoArray Project with 13,793 NSCLC cases and 14,027 controls).We performed genome-wide scan on the NJMU GSA Project in the present study,and the other GWASs were existing datasets.The genome-wide genotype data of each GWAS in the NJMU GSA Project had undergone standardized quality controls.Briefly,we excluded samples with call rate<95%or sex discrepancy,or samples with extreme heterozygosity;when a pair of samples were duplicates or first degree relatives,we excluded the sample with lower call rate;population outliers were also excluded.For genetic variants,we excluded duplicate markers or variants with call rate<95%;variants that failed cluster inspection,or not in Hardy-Weinberg equilibrium,or with minor allele frequencies(MAFs)<0.1%were excluded.Detailed quality control processes of existing GWASs had been described in papers published previously.The qualified genotypes for each chromosome were phased with SHAPEIT,and imputation was performed using IMPUTE2 based on haplotypes derived from the 1000 Genomes Project(the Phase III integrated variant set release).For each GWAS dataset,per-allele odds ratios(ORs)and standard errors(SEs)were calculated using logistic regression.Age,sex,pack-year or smoking status,and the first ten principal components(PCs)were adjusted in the NJMU GSA Project,the NJMU GWAS,and the NJMU OncoArray GWAS.While for the TRICL-ILCCO OncoArray Project,we adjusted for age,sex,and the first three PCs.A meta-analysis of fixed-effects model was performed using METAL software to combine individual association estimate from each GWAS.We performed meta-analysis on six GWASs of Chinese populations,as well as trans-ancestry meta-analysis on all GWAS datasets.We set P≤5×10-8 as the genome-wide significance threshold.The heterogeneity in genetic effects across studies was assessed by using Cochran’s Q test.We only included variants with at least two-thirds of the samples contributing to the meta-analysis.Variants with I2≥75%or P value for Cochran’s Q statistic≤1×10-4were excluded from the analysis.We conducted genome-wide Meta-analysis for NSCLC,lung adenocarcinoma(LUAD),and lung squamous cell carcinoma(LUSC),respectively.Heterogeneity of the effect estimates for risk SNPs across Chinese and European populations were assessed using the Cochran’s Q test.We also set I2≥75%or P value for Cochran’s Q statistic≤1×10-4 as indicating a high degree of heterogeneity.To identify interactions between risk SNPs of NSCLC and smoking pack-years,an interaction analysis was performed in Chinese and European populations,respectively,by including the interaction term in the logistic regression.Finally,we estimated the heritability that can be explained by lung cancer risk loci in Chinese populations using the liability threshold method.We built polygenic risk score(PRS)based on these risk SNPs.The associations of PRS,smoking pack years,and their combinations with risk of NSCLC were analyzed.We also built risk prediction models that included age,gender,smoking pack years,and PRS to investigate the potential utility of PRS in risk stratification of lung cancer.Results:The present study identified 19 risk loci in NSCLC,LUAD,and LUSC,among which six were novel loci.The six novel susceptibility loci included three loci for overall NSCLC risk:rs3769821 in 2q33.1(OR=1.08,P=4.45×10-8),rs2293607 in3q26.2(OR=1.10,P=1.82×10-10),and rs1200399 in 14q13.1(OR=1.11,P=3.05×10-9);two for LUAD:rs17038564 in 2p14(OR=1.15,P=1.87×10-8),and rs35201538 in 9p13.3(OR=1.10,P=1.99×10-8);and one for LUSC:rs4573350 in 9q33.2(OR=1.13,P=3.23×10-9).Among the 13 risk loci that had been reported by lung cancer GWASs,8p12 and 11q23.3 were reported as lung cancer risk loci in European populations,and they also reached genome-wide significance according to the association results of NSCLC in Chinese populations.We compared effect estimates of the 19 risk loci between Chinese and European populations.Seven out of the 19 risk loci(i.e.,3q28,5p15.33,6p21.1,8p12,9q33.2,15q25.1,and 19q13.2) showed high degree of heterogeneity(I 2≥75%or P value≤1.0×10-4)across Chinese and European populations.The gene-environment interaction analysis identified two pairs of significant interactions:the interaction between rs62332588(5p15.33)and smoking pack-years in Chinese populations(P=7.44×10-7),and the interaction between rs55781567(15q25.1)and smoking pack-years in European populations(P=2.26×10-8).From 81 lung cancer risk SNPs that were either reported previously or identified in the present study,we selected 19 independent risk SNPs of NSCLC in Chinese populations.The heritability that can be explained by the 19 risk SNPs in Chinese populations was 1.00%.In addition,we calculated the PRS based on the 19risk SNPs for each individual,and association analysis found that the PRS was significantly associated with increased risk of NSCLC.Compared to participants with PRS in the lowest decile,the risk of developing NSCLC increased by 0.57 times for those with PRS in the fifth decile(OR=1.57,P=2.48×10-15);the risk of developing NSCLC increased by 1.79 times for participants with PRS in the top decile(OR=2.79,P=7.66×10-70).Within groups of never smokers,light smokers,or heavy smokers,PRS was associated with increased risk of NSCLC.We constructed risk prediction models that included age,gender,smoking pack-years,and PRS.The area under the curve(AUC)of the“age+gender+smoking pack-years”model was 0.622.And the AUC of the“age+gender+smoking pack-years+PRS”model was 0.649,whose discrimination ability had significantly improved(P=2.75×10-28).Conclusions:The present study identified six novel susceptibility loci for NSCLC,which contributed to our understanding of the genetic architecture for NSCLC. Moreover,PRS is an independent risk factor for NSCLC,and can improve the risk stratification of NSCLC.These results will facilitate individualized prevention of lung cancer in the future. |