Font Size: a A A

Research On P2P Loan Defaults Prediction Based On CatBoost Stacking Method

Posted on:2024-06-29Degree:MasterType:Thesis
Country:ChinaCandidate:W B LaiFull Text:PDF
GTID:2568307091991319Subject:Applied Statistics
Abstract/Summary:
In recent years,many traditional financial enterprises have accelerated their digital transformation,gradually shifting from a heavy front office over the back office to a light front office and a heavy back office,and the tide of online traditional offline loans is unstoppable.In view of the characteristics of short term,small amount,high frequency and urgent demand for small and micro credit and personal credit business,in addition to the existing P2 P online loan business,online loan products for individual consumers and small and micro enterprises launched by banks continue to emerge,which promotes the healthy development of the online loan market.At the same time,the development of big data and credit reporting system has made credit information more digital,structured and standardized,and the actual cost of obtaining corresponding data by financial institutions has been greatly reduced.Therefore,how to timely and efficiently use and analyze relevant data,discover high-risk behaviors such as fraud,gang fraud,default and bad debts that may occur among applicant users,and ensure the safety of funds is an urgent problem to be solved in the healthy development of Internet financial consumer credit.Therefore,based on the existing relevant research literature at home and abroad,this thesis summarizes the research results of relevant scholars and points out the shortcomings of existing relevant research results.Based on this,a credit risk control model based on the CatBoost model is first proposed,and then the stacking model is used to fuse ideas and the performance differences of each model,and the CatBoost model is combined with logistic regression,XGBoost and Light GBM to form the CatBoost-Stacking model,which takes logistic regression,XGBoost and Light GBM as the base model,and CatBoost as the metamodel.The classification results of the base model and the dataset are used as the input of the metamodel,and the classification results of the metamodel are used as the final output results,which avoids the problem of lack of attention to some features and samples under a single model,and effectively improves the prediction accuracy of the model.The dataset selected in this thesis is the desensitization transaction data of 2019 and 2020 published by the Lending Club platform,and the dataset is first descriptively analyzed and preprocessed,and then the dataset is feature-engineered.In order to test the stability of the model,the dataset is randomly split into two parts,the training set and the test set,and the test set is used as the final evaluation criterion of the model.In order to evaluate the advantages and disadvantages of the CatBoost model,comparing it with logistic regression,GBDT,XGBoost and Light GBM,it is found that CatBoost is the best in both accuracy and AUC values,where the accuracy is 0.023 higher than Light GBM,0.009 higher than XGBoost,and the AUC value is 0.005 higher than Light GBM and 0.007 higher than XGBoost.For the CatBoost-Stacking model,compared with the single CatBoost model,the AUC value and KS value were greatly improved,with the AUC value being 0.006 higher and the KS value being 0.012 higher.Therefore,both the CatBoost model and the CatBoost-Stacking model can effectively improve the classification ability and accuracy of prediction and evaluation,and provide a new solution for the research of personal credit assessment risk problem.Finally,the shapley additive explanation is used to evaluate the impact of different characteristics on default risk,and suggestions are made to financial institutions: the model team needs to keep pace with the times,continuously monitor the model performance from various aspects such as classification quality,stability,and operational efficiency,and iterate the model in a timely manner;The model team should work closely with the local marketing department to model loan varieties with different loan amounts,regions,and credit ratings;Establish a risk control model system for different stages of pre-loan,loan and post-loan to learn from each other’s strengths;Third-party data introduced by financial institutions should balance the availability and authenticity,and can be verified by multiple parties through other information to avoid outliers affecting model quality.
Keywords/Search Tags:Credit risk, Ensemble Learning, CatBoost, SHAP
Related items