Font Size: a A A

A Study Of Bayesian Classifiers For Categorical Variable

Posted on:2012-12-29Degree:DoctorType:Dissertation
Country:ChinaCandidate:G DuFull Text:PDF
GTID:1100330338490588Subject:Probability theory and mathematical statistics
Abstract/Summary:
Classification , which is broadly applied in Bioinformatics, statistical physics, fi-nance, industrial manufacture, quality control and so on, is one of the core tasks inStatistical Inference. Through continual endeavor, researchers have proposed plentyof methods for classification, such as Fisher Discrimination Analysis, Logistic Regres-sion, lasso, neutral networks, SVM. HoIver, due to the rapid development of scienceand technology, people encounter many new emerging problems in practice, which arenew challenges to statisticians.For example, in Bioinformatics, researchers usually aim at the identification ofgenes associated with some specific disease, and predict the present of this diseaseaccording to these genes associated. The difficulties is that the amount of candidategenes that may be related to the disease is far greater than that of patients available.In statistical language, the problem is how to effectively classify the high dimensionalcategorical data when the sample size is much smaller than the number of covariates.High dimensional data, especially for the large p, small n data, is not rare in prac-tice. My thesis addresses the classification problem of the categorical data. We considerboth situations that the number of covariates is smaller than sample size and verse vice.During our research, I propose two novel Bayesian classifiers: SPAN-2 and STAN andfurther generalize them to GSPAN-2 and GSTAN to solve the possible distraction ofinteracted noisy covariates.SPAN (Yuan [2009]) generalizes the idea of BEAM (Zhang and Liu [2007]). AndI propose SPAN-2 that resolve deficiencies of SPAN. SPAN-2 adopts a new variationof Metropolis algorithm in order to preventing SPAN from trapping in local mode.Therefore, SPAN improve the MCMC efficiency of SPAN. Simulation study illustratesthat SPAN-2 outperforms SPAN in terms of classification accuracy.Then, I innovationally combine the idea of partition of covariates in BEAM withthe idea of netting covariates in a tree in TAN (Friedman et al. [1997]), and proposea new classifier: STAN. We employ MTM (multiple-try Metropolis) technique to con-struct STAN, while TAN uses a exhaustive search. Hence, albeit STAN is more com-plicated than TAN, both algorithms share the same computation complexity O(L2·N),where L is the number of covariates and N is sample size. STAN partitions all co- variates into three non-overlapping groups, all noisy covariate are partitioned in group1, and all informative covariates are partitioned into two groups based on their cor-relations. Intuitively, group 2 contains all informative covariates that in?uence classvariable independently, and group 3 contains all informative covariates that in?uenceclass variable jointly. For group 3, I use a Bayesian network to depict their interactions(i.e. correlation structure). This modeling of covariates endows STAN the ability tosimultaneously achieve variable selection and identification of variable interaction.In simulation studies and real data analysis, STAN is competing regarding to clas-sification accuracy, particularly, STAN outperforms other classifiers when Signal-to-Noise ratio is low. Moreover, STAN is able to capture informative covariates as Ill astheir interactions. As a result, STAN exhibits sound robustness: for three different situ-ations: 1. L=50,N=400, 2.L=500,N=400, 3.L=2000,N=400, the accuracies of STAN isalmost the same. In contrast, the accuracies of other classifiers depreciates in varyingdegrees as the number of covariates increases.Especially for simulation study 2, I simulated the informative covariates that in-teract with each other without marginal effects on class variable. Our STAN effectivelyidentifies these informative covariates while other methods fail. As a result, STAN'sclassification accuracy is far higher than that of others.Finally, I further generalize SPAN-2 and STAN. Most previous classifiers doesn'ttake account of the identification of interacted covariates, which might be wrongly rec-ognized as informative and, in turn, result in lowing classification accuracy or increas-ing variance of the model. To tackling this problem, I further split noisy covariatesinto two subgroups, one contains all independent noisy covariates, another one con-tains correlated noisy covariates. In summary, all covariates are therefore partitionedinto 4 groups: two groups for noisy covariates, two for informative covariates. Basedon this new partition idea, I generalize SPAN-2 and STAN to GSPAN-2 and GSTANrespectively. GSPAN-2 and GSTAN effectively solve the problem that noisy covariateare wrongly grouped as the simulation study illustrates. Hence, GSPAN-2 and GSTANown better classification capabilities and noise resistance characteristics.
Keywords/Search Tags:Bayes, MCMC, Classification, Variable selection, Variable interaction
Related items