Lung cancer predictive factor analysis:construction of a prediction model based on machine learning
Wan Chaopeng
Peng Changmei
Mo Qihua
Liu Lili
Lin Kun
Qiu Weili
Lin Kai
Abstract:Objective To explore the potential risk factors for the occurrence of lung cancer and use machine learning(ML)al-gorithms to establish a lung cancer risk prediction model,providing scientific evidence for the prevention and treatment of lung cancer and achieving precise prevention.Methods A total of 1,066 newly diagnosed lung cancer patients,confirmed by pathol-ogy,admitted to a hospital in Guangdong from January 2013 to December 2023,were selected as the lung cancer group.Simul-taneously,1,066 individuals with no lung cancer or other cancers,selected from a health check-up center at a 1∶1 ratio,were included as the control group.Data,including demographic information and clinical data,were collected from both groups.The dataset was randomly divided into a training set(70%)and a testing set(30%).The most meaningful features were se-lected for the model using the Least Absolute Shrinkage and Selection Operator(LASSO)algorithm.Eight common ML algo-rithms were employed to establish prediction models:Logistic Regression(LR),Random Forest(RF),Multilayer Perceptron(MLP),Support Vector Machine(SVM),K-Nearest Neighbors(KNN),Light Gradient Boosting Machine(LightGBM),Ex-treme Gradient Boosting(XGBoost),and Decision Tree(DT).Bayesian optimization was applied to each model,and five-fold cross-validation was performed to test the best hyperparameters for each model.The performance of these models was evalua-ted using Receiver Operating Characteristic(ROC)curves and Decision Curve Analysis(DCA).The SHAP algorithm was used to interpret the best-performing model,enhancing the visualization of important features.A new model was built based on Stacking ensemble learning to further promote model performance.Results The LASSO algorithm identified important varia-bles,including age,smoking status,body mass index(BMI),fruit and vegetable consumption,exercise duration,education level,diabetes history,family history of cancer,family history of lung cancer,chronic bronchitis,emphysema history,chronic obstructive pulmonary disease(COPD)history,long-term exposure to coal smoke,long-term exposure to cooking oil smoke,and long-term exposure to smoke from firewood.In the prediction model,the best performance was observed in the LightGBM model,which had an area under the curve(AUC)of 0.961 8,accuracy of 0.893 8,precision of 0.894 0,sensitivity of 0.880 9,speci-ficity of 0.906 5,recall of 0.893 8,and F1 score of 0.893 7,showing good overall predictive performance.In the DCA curve,the net benefit of the LightGBM model exceeded that of the other models.SHAP analysis indicated that BMI,age,and emphyse-ma history contributed most to lung cancer risk.The new model,based on Stacking ensemble learning,showed an accuracy higher than that most of the single models,with the SVM model as the meta-learner achieving the highest accuracy of 89.06%and an AUC of 0.915 1.Conclusion This study successfully constructed a lung cancer risk prediction model based on LightG-BM and Stacking ensemble learning,confirming multiple predictive factors for lung cancer risk.In clinical practice,the use of these predictive factors can quickly identify high-risk patients,providing an effective tool for early identification and prevention and contributing to more effective prevention and treatment measures for lung cancer.
Keywords:Machine learningLung cancerPredictive factorsPrediction modelStacking ensemble
Publication Date:2025-05-20
Online Publishing Date:2025-09-04(First online date of this platform, not the publication date of the document)
Pages:7( 50-56 )
