Font Size: a A A

Model Structure Learning For High Dimensional Complex Data

Posted on:2024-01-24Degree:DoctorType:Dissertation
Country:ChinaCandidate:Y H GeFull Text:PDF
GTID:1520307307495384Subject:Applied Statistics
Abstract/Summary:
High-dimensional data provide rich descriptions of biomedical,economic and financial problems but also pose challenges in statistical modeling,computation and prediction.In data analysis,summarizing useful information and discovering the low-dimensional structure from the huge amount of high-dimensional data are important for enhancing the explanatory power,reducing computational cost and improving subsequent prediction of statistical analysis.Identifying the low-dimensional structure of data can often be summarized as a statistical task named as model structure learning.Specifically,variable selection is an important part of model structure learning.The ideas and methods in variable selection are widely applied to other tasks in the model structure learning,such as parametric/nonparametric structure identification and interaction effect identification.The purpose of variable selection is to identify a few covariates associated with the response variable from a large number of covariates,and to avoid model overfitting and prediction failure due to redundant variables.Moreover,under the ultra-high dimensional settings where the dimension of variables grows with the sample size at some exponential rate,the performance of variable selection based on regularization regression deteriorates due to certain severe spurious correlations among the covariates.Some studies suggest that it can improve the efficiency of statistical analysis by identifying some covariates with model-free style feature screening methods and then implementing variable selection with those identified covariates.The problem of variable selection and model structure learning has been studied at length in the field of statistics and machine learning,but most of the studies depends on strong assumptions about the model specifications and the signal strength of the important variable.This limits the generality and reliability of variable selection and model identification methods.The interpretability and reproducibility of variable selection has also been a widespread concern in applications.In recent years,some scholars have started to work on the quantitative evaluation of variable selection,and the idea of multiple testing has been introduced in the variable selection problem to control the false discovery rate.These developments have also deepened our thinking on the tasks of variable selection and model structure learning.In Chapter 1,we first introduce the feature screening and model structure learning methods,and give a review of its recent advances.In Chapter 2,we consider the problem of feature ranking in the context of identification of genomic markers.A number of statistical methods have been developed to search for genomic markers associated with the development and progression of diseases.Among them,feature ranking plays a vital role due to its intuitive formulation and computational efficiency.However,most of the existing methods are based on the marginal importance of molecular predictors and share the limitation that the dependence(network)structures among predictors are not well accommodated.In this chapter,we propose a structured feature ranking method for identifying genomic markers,where such network structures are effectively accommodated using Laplacian regularization.The proposed method investigates multiple network scenarios,where the networks can be known a priori or data-dependently estimated.In addition,we rigorously explore the noise and randomness in the networks and control their impacts with the proper selection of tuning parameters.These characteristics make the proposed method enjoy especially broad applicability.Theoretical result of our proposal is rigorously established,and the statistical guarantee is also given for networks with general Kirchnoff matrices under mild conditions.Extensive simulations and analysis of The Cancer Genome Atlas melanoma data demonstrate the improvement of finite sample performance and practical usefulness of the proposed method.In Chapter 3,a novel false discovery rate(FDR)control method for Cox model is proposed,which combines uneven data splitting,symmetric-based statistic and decorrelated estimators.The key step is to construct a sequence of ranking statistics based on two independent regression coefficients via uneven sample splitting.FDR control is achieved by chooses a data-driven threshold along the ranking.Thanks to the benefits of uneven data split and decorrelated method,our method can provide more robust variable selection results than existing methods in finite sample analysis.We establish the asymptotic theory on FDR control of the proposed approach at any designated level.Extensive simulation studies and an empirical application on a P2P loan data confirm the robustness of the proposed method in FDR control,and show that it often achieves higher power among competitors.In Chapter 4,unlike most of those existing methods that focus on some specific settings under certain model assumptions,this work proposes a general and novel framework for recovering true structures of target functions by using unstructured M-estimation in a Reproducing Kernel Hilbert Space(RKHS).The proposed framework is inspired by the fact that gradient functions can be employed as a valid tool to learn underlying structures,including sparse learning,interaction selection and model identification,and it is easy to implement by taking advantage of the nice properties of the RKHS.More importantly,it admits a wide range of loss functions,and thus includes many commonly used methods,such as mean regression,quantile regression,likelihood-based classification,and margin-based classification,which is also computationally efficient by solving convex optimization tasks.The asymptotic results of the proposed framework are established within a rich family of loss functions without any explicit model specifications.The superior performance of the proposed framework is also demonstrated by a variety of simulated examples and a real case study.The dissertation focuses on the problem of model structure learning for high-dimensional data.All the methodology proposed in this dissertation is implemented with R and available upon requirement for the convenience of scientific use.
Keywords/Search Tags:High Dimensional Data Analysis, Feature Ranking, Variable Selection, Interaction pursuit, Multiple Testing
Related items