Stochastic stepwise ensembles for variable selection

Posted on:2010-04-10

Degree:M.Math

Type:Thesis

University:University of Waterloo (Canada)

Candidate:Xin, Lu

Full Text:PDF

GTID:2440390002486748

Subject:Statistics

Abstract/Summary:

Ensembles methods such as AdaBoost, Bagging and Random Forest have attracted much attention in the statistical learning community in the last 15 years. Zhu and Chipman (2006) proposed the idea of using ensembles for variable selection. Their implementation used a parallel genetic algorithm (PGA). In this thesis, I propose a stochastic stepwise ensemble for variable selection, which improves upon PGA.;Instead of adding or deleting one variable at a time, Stochastic Stepwise Algorithm (STST) adds or deletes a group of variables at a time, where the group size is randomly decided. In traditional stepwise, the group size is one and each candidate variable is assessed. When the group size is larger than one, as is often the case for STST, the total number of variable groups can be quite large. Instead of evaluating all possible groups, only a few randomly selected groups are assessed and the best one is chosen.;From a methodological point of view, the improvement of STST ensemble over PGA is due to the use of a more structured way to construct the ensemble; this allows us to better control over the strength-diversity tradeoff established by Breiman (2001). In fact, there is no mechanism to control this fundamental tradeoff in PGA. Empirically, the improvement is most prominent when a true variable in the model has a relatively small coefficient (relative to other true variables). I show empirically that PGA has a much higher probability of missing that variable.;Traditional stepwise regression (Efroymson 1960) combines forward and backward selection. One step of forward selection is followed by one step of backward selection. In the forward step, each variable other than those already included is added to the current model, one at a time, and the one that can best improve the objective function is retained. In the backward step, each variable already included is deleted from the current model, one at a time, and the one that can best improve the objective function is discarded. The algorithm continues until no improvement can be made by either the forward or the backward step.

Keywords/Search Tags:

Variable, Stochastic stepwise, Ensemble, Selection, PGA, Backward, Forward

Related items

1	High Order Numerical Methods And Error Estimates For Solving Decoupled And Weakly Coupled Forward-Backward Stochastic Differential Equations
2	Some Optimal Control And Differential Game Problems In Forward-Backward Stochastic Systems
3	The Maximum Principle For Optimal Control Problem Of Mean-Field Forward-Backward Stochastic System With State Constrained
4	Exploration And Generalization Of Several Stepwise Variable Selection Algorithms
5	Optimal Control Problems Of Some Forward-Backward Stochastic Pantograph Systems
6	High Dimensional Backward Stochastic Differential Equations, Forward-Backward Stochastic Differential Equations And Their Applications
7	Mean-Field Forward-Backward Stochastic Differential Equations And The Related Questions
8	Rank-based Forward Backward Stochastic Differential Equations And Nonlinear Expectation
9	Optimal Control And Differential Game Of Partial Information Forward-Backward Stochastic Systems
10	H₂/H_∞ Control Of Forward And Backward Stochastic Systems