| Bruton’s Tyrosine Kinase(BTK)is a non-receptor tyrosine kinase belonging to the TEC family,which is crucial for the proliferation and survival of malignant tumors derived from B cells.Therefore,it is considered an important anti-tumor target.However,all approved BTK inhibitors are first-generation inhibitors that covalently bind to BTKCys481.This dependence on BTKCys481 can lead to acquired resistance of inhibitors,weakening their inhibitory effects.Thus,it is necessary to develop second-generation non-covalent BTK inhibitors and devising effective strategies to overcome resistance.This study aims to construct a comprehensive human BTK inhibitor database,predict their biological activity using computer-aided methods,and investigate the relationship between the structure and activities.Through these efforts,we can better understand the mechanism of BTK inhibitors and provide essential theoretical guidance for the further development of efficient anti-tumor drugs.The main work of this study includes:1.Build qualitative classification models of non-covalent BTK inhibitors by machine learning methodsWe first collected BTK inhibitor data from the Reaxys and Ch EMBL databases and conducted data preprocessing.To guide the design of next-generation non-covalent BTK inhibitors with ideal activity,we selectively removed compounds that may form covalent bonds with the residue C481during data preprocessing,resulting in a dataset of 3,895 inhibitors.To characterize these inhibitors,we used MACCS fingerprints and Morgan fingerprints for feature extraction,and built 16 classification models with four traditional machine learning algorithms:Decision Tree(DT),Random Forest(RF),Support Vector Machine(SVM),and Extreme Gradient Boosting(XGBoost).Additionally,two deep learning methods were employed,including Deep Neural Networks(DNN)and Recurrent Neural Networks(RNN).The Model D_4 using XGBoost and MACCS fingerprints shows the best classification performance.With an prediction accuracy of 94.1%and a Matthews Correlation Coefficient of 0.75 on the test set.2.Bulid clustering models of molecular structure of non-covalent BTK inhibitorsTo perform a more in-depth analysis of the 3,895 BTK inhibitors,we employed the K-means clustering algorithm,dividing them into 9 subclasses.We validated the reliability of the clustering model by cross-validating the results of the SHAP value hierarchical clustering method and K-means dimensionality reduction visualization.We found two subclasses that needed further subdivision into smaller subcategories.To better understand the model’s predictions,we used the SHAP method to decompose the predicted values,obtaining the contributions of each feature to the prediction results.By inputting each subclass obtained from K-means clustering into the SHAP for analysis of the classification models.This study revealed the importance of various molecular scaffolds in different subclasses.Comparing these results with existing crystal structure data confirmed the reliability and specificity of SHAP analysis.This also further verified the accuracy of the qualitative model,proving its value in guiding the design of next-generation BTK inhibitors with ideal activity.3.Research on the quantitative regression model of BTK inhibitorsWe also constructed several quantitative structure-activity relationship models on BTK inhibitors using machine learning algorithms.During data collection and preprocessing,we selected inhibitors with determined IC50 values as the data source,using the negative logarithm of IC50 as the activity index.In the modeling process,we used two dataset partitioning methods,three types of physicochemical property descriptors,and three machine learning algorithms were used for modeling.The results showed that the model J using SOM partitioning and Rdkit descriptors generally outperformed other datasets on the test set,with Model J_6 using Extreme Gradient Boosting achieving the best performance(R2=0.83,RMSE=0.36).Finally,we used the SHAP method to explain the impact of each descriptor on the activity in the BTK inhibitor activity prediction model and analyzed their physical and chemical significance.In this study,we constructed a BTK inhibitor database and developed bioactivity prediction models using various machine learning algorithms.Through clustering studies and SHAP analysis,the interpretability of the models was examined.Ultimately,we established quantitative regression models to investigate the influence of physicochemical properties on BTK inhibitor activity,providing theoretical support for the design of the next generation of non-covalent BTK inhibitors. |