UOJM Volume 16, Issue 1

RESEARCH

Using AI to Predict Immunotherapy Response in Non-Small Cell Lung Cancer Patients

Kyle Cheng1, Aaron Sethi1
1 Carleton Univeristy, Ottawa, ON, Canada

University of Ottawa Journal of Medicine, Volume 16, Issue 1, July 2026, pg. 41-48, https://doi.org/10.18192/UOJM.V16i1.7562
First published online June 28, 2026.

Keywords: non-small cell lung cancer (NSCLC), machine learning, immunotherapy, logistic regression, Immune Checkpoint Inhibitors (ICI).


Abstract

Over the past decade, immunotherapy has emerged as one of the most promising treatments for cancer. Specifically, immune checkpoint inhibitors (ICIs) have revolutionized treatment for non-small cell lung cancer (NSCLC), yet only 20–30% of patients experience an objective response. Reliable biomarkers include programmed death-ligand 1 (PD-L1) tumour proportion score and tumour mutational burden (TMB), though they are not universally predictive alone. This study aims to improve patient selection by developing a logistic regression machine learning model for ICI response prediction in NSCLC patients. We analyzed a public dataset of 224 ICI-treated NSCLC patients from Memorial Sloan Kettering (MSK) Cancer Center, incorporating clinical, pathologic, and genomic data. Baseline features such as age, PD-L1 tumour proportion score, TMB, neutrophil-to-lymphocyte ratio (NLR), pack-year smoking history, among others were evaluated. Exploratory data analysis was followed by model training and optimized with Lasso/L1 regularization. Performance was assessed using ROC-AUC, F1 scores, and 5-fold cross-validation. The final model achieved an ROC-AUC of 0.84, 80% accuracy, an F1 score of 0.67, and recall of 1.0. Responder precision reached 50%, compared to non-model guided rates of 20–30%. Previous approaches such as deep learning, multimodal feature fusion and other black-box approaches are limited by their complexity and lack of interpretability, posing significant challenges for real-world clinical application. In contrast, our white-box model employs a prospective, robust, and efficient design that re-trains on new, real-time data, spares non-responders from unnecessary costs, and advances personalized immunotherapy research.

Résumé

Au cours de la dernière décennie, l’immunothérapie est devenue l’un des traitements les plus prometteurs contre le cancer. Plus précisément, les inhibiteurs de points de contrôle immunitaire (ICI) ont révolutionné le traitement du cancer du poumon non à petites cellules (CPNPC), mais seulement 20 à 30 % des patients présentent une réponse objective. Les biomarqueurs fiables comprennent le score de proportion tumorale du ligand 1 de mort programmée (PD-L1) et la charge mutationnelle tumorale (TMB), bien qu’ils ne soient pas universellement prédictifs à eux seuls. Cette étude vise à améliorer la sélection des patients en développant un modèle d’apprentissage automatique basé sur la régression logistique pour prédire la réponse aux ICI chez les patients atteints de CPNPC. Nous avons analysé un ensemble de données publiques portant sur 224 patients atteints de CPNPC traités par ICI au Memorial Sloan Kettering (MSK) Cancer Center, en intégrant des données cliniques, pathologiques et génomiques. Des paramètres initiaux tels que l’âge, le score de proportion tumorale de PD-L1, la TMB, le rapport neutrophiles/lymphocytes (NLR) et l’antécédent tabagique en paquets-années, entre autres, ont été évaluées. Une analyse exploratoire des données a été suivie de l’entraînement du modèle, optimisé par une régularisation Lasso (L1). Les performances ont été évaluées à l’aide de l’aire sous la courbe ROC (ROC-AUC), des scores F1 et d’une validation croisée en 5 plis. Le modèle final a atteint une ROC-AUC de 0,84, une exactitude de 80 %, un score F1 de 0,67 et un rappel de 1,0. La précision chez les patients répondeurs a atteint 50 %, comparativement à un taux de 20 à 30 % en l’absence de guidage par modèle. Les approches précédentes, telles que l’apprentissage profond, la fusion de caractéristiques multimodales et d’autres méthodes de type « boîte noire », sont limitées par leur complexité et à leur manque d’interprétabilité, ce qui pose des défis importants pour une application clinique en pratique réelle. À l’inverse, notre modèle de type « boîte blanche » repose sur une conception prospective, robuste et efficace qui permet de se réentraîner sur de nouvelles données en temps réel, évite aux non-répondeurs des coûts inutiles et avancer la recherche en immunothérapie personnalisée.



Introduction

Immunotherapy, a cancer treatment that leverages a patient's own immune system to fight cancer cells, has transformed current therapies over the last decade1. One of the most adopted forms is immune checkpoint inhibitors (ICIs), which target regulatory molecules such as PD-1, PD-L1, and CTLA-4 to prevent tumour-induced immune suppression2,3. Non-small cell lung cancer (NSCLC), the most common type of lung cancer, accounting for 80% to 85% of all lung cancer cases,4 has been the primary focus of ICI therapies2,3. Despite their clinical promise and potential, ICI treatments on NSCLC patients achieve an objective response rate of 20-30%5. These findings reveal the limited clinical benefit this treatment currently delivers, highlighting a need for optimized patient selection to maximize accurate outcomes.

Current clinical research primarily focuses on using biomarkers such as programmed death-ligand 1 (PD-L1) and tumour mutational burden (TMB) to guide patient selection for ICI therapy6,7. PD-L1 tumor proportion score measures the percentage of cells that express PD-L1, while TMB quantifies the number of genetic mutations in the tumor as the number of mutations per megabase of DNA. While these biomarkers offer predictive value, they are inconsistent and vary across individuals due to factors such as tumour heterogeneity, dynamic immune interactions, and spatial variability in expression6,7. For these reasons, while TMB and PD-L1 remain valuable biomarkers, they are not sufficient to fully account for patient response variability, requiring a need for more robust and individualized tools that enhance decision-making.

Recent advancements in machine learning (ML) and artificial intelligence (AI) have accelerated progress toward more personalized medicines8,9. ML models have shown a high degree of effectiveness when trained on high-dimensional clinical and molecular data,8,9 often able to spot unique non-linear patterns and relationships that are not captured by traditional statistical methods. In the case of ICI treatment of NSCLC, these approaches can be used to predict patient responses with the goal of improving patient selection10,11.

This study uses a supervised machine learning approach by using logistic regression to predict immunotherapy response in NSCLC patients. Previous machine learning models, such as deep learning models used by Rakaee et al., or multimodal feature fusion used by Vanguri et al., are considerably more complex, posing significant challenges to explainability in clinical settings12,13. In contrast, logistic regression provides greater transparency and reliability, making it a more desirable choice for clinical application14.  By using a public dataset of 224 ICI-treated NSCLC patients from the Memorial Sloan Kettering (MSK) Cancer Center15, we evaluated the predictive power of commonly available clinical features, including PD-L1, TMB, age, neutrophil-to-lymphocyte ratio (NLR), smoking history (pack-year), and microsatellite instability (MSI). Specifically, pack-year smoking history refers to the number of cigarettes a patient smoked per day for a given number of years. 

The aim of this study is to develop a model that can act as a clinical decision support tool for oncologists by enhancing precision in patient selection for immunotherapy16. As a result, this study poses the question: Can machine learning be used to predict immunotherapy response in NSCLC patients, and can this approach overcome the limitations of existing therapies?

Methods

In this study, we analyzed 224 ICI-treated NSCLC patients from the MSK Cancer Center.15 The MSK-IMPACT dataset was used for this study and it was chosen over other datasets because it is one of the largest, prospective cancer sequencing datasets available, making it ideal for broad genotypic and phenotypic analysis.15,17 Patients of any age with NSCLC tumor/normal pairs treated at MSKCC with anti-PD-(L)1 based therapy were included in this data. Patients with missing data were excluded for this study as they lacked the necessary information for this analysis. The MSK-IMPACT Non-Small Cell Lung Cancer dataset operates under Memorial Sloan Kettering Cancer Center IRB Protocol 12-245, with data gathered via patient-informed consent protocols. The therapeutic response in each patient was evaluated according to the Response Evaluation Criteria in Solid Tumours (RECIST v1.1)18, which allowed for subdivision into two groups: responders and non-responders. RECIST is the standard guideline to objectively measure the response to treatment. Responders are characterized by a complete response (CR), where all target lesions disappear, or a partial response (PR), defined as a 30% decrease in the sum of the longest diameters of target lesions compared to baseline. Non-responders are classified as either progressive disease (PD), a 20% increase in sum of the longest diameters of target lesions or stable disease (SD), indicating no significant shrinkage18. To ensure sufficient data quality and clinical relevance, only the patients diagnosed with NSCLC, without prior ICI treatment and having six to eight baseline characteristics, including PD-L1 expression and TMB, were included. Patients were excluded if they had active autoimmune disease, ongoing immunosuppressive therapy, or were missing any key baseline characteristic data.

Features used to train the model were easily accessible clinical and molecular variables: age, PD-L1 TPS, Tumour Mutational Burden (TMB), Neutrophil Lymphocyte Ratio (NLR), pack-year smoking history, microsatellite instability, albumin, and fraction of genome altered (FGA). These variables capture complementary aspects of tumour biology, immune activity, and overall patient health, making them well suited to guide predictive modelling. The target variable was the patient response to ICI therapy.

Preprocessing was done through exploratory data analysis (EDA) to examine distributions, correlations, and identify potential outliers. Numerical values were standardized using Z-score normalization. Missing target values resulted in exclusion of that data point, and data with missing features were excluded by listwise removal. There was a class imbalance between responders and non-responders for which we employed synthetic oversampling through SMOTE (Synthetic Minority Over-sampling Technique)19 on the training data to help model learning. This addresses the imbalance by generating synthetic samples for the minority class, thereby balancing class representation and reducing model bias toward the majority class.

This study chose logistic regression due to its increased interpretability and clinical transparency and compared L1 and L2 regularization to reduce overfitting and address multicollinearity. The data was divided into 80% training and 20% testing (Figure 1). Hyperparameters were optimized using grid search and 5-fold cross-validation on training data (Figure 2). The optimal probability threshold of 0.42 for classification was determined using the Youden J statistic that was 0.74.

Figure 1. Flowchart showing model testing and training split and optimization process

Figure 2. The distribution of the Best Overall Response (BOR) using Seaborn and Matplotlib. (1 = Responder, 0 = Non-responder)

The performance of the model was evaluated on the test set using ROC-AUC, accuracy, precision, recall, F1 score, and confusion matrix. The analysis was done in Python (3.9.6) using libraries such as pandas, scikit-learn, seaborn, and matplotlib.20,21,22,23,24

Results

A total of 224 NSCLC patients treated with ICI were analyzed. Of these, 67 responded to therapy and 157 did not respond, based on RECIST v1.1 evaluation criteria18 (Figure 3). Exploratory data analysis (EDA) identified that TMB and PD-L1 levels were significantly higher in responders than non-responders, suggesting that these variables may be associated with improved predictive performance (Figure 4). 

Figure 3. Relationship of PD-L1 TPS and TMB (n = 223 patients) between responders (PR/CR) and nonresponders (SD/PD). The box-and-whisker bars include the mean, the IQR (25–75%) and the minimum and maximum (up to 1.5 × IQR) as whiskers.

Figure 4. Heatmap created using Seaborn. Displays the Spearman correlation coefficients between features. Best overall response (BOR) is the target variable.

A Spearman correlation matrix was made to analyze the relationships between features (Figure 5). The feature coefficient analysis showed that pack-year smoking history, TMB, and PD-L1 expression were the most influential features in the dataset, consistent with established clinical data (Figure 5). While there were some correlations between the clinical variables, there was no strong collinearity, which supported the inclusion of these variables into our multivariable model.

Figure 5. Applied 5-fold cross-validation of Grid Search to search for optimal hyperparameters (Lasso and Ridge regularization), optimizing ROC-AUC.

On the testing set, the model achieved an ROC-AUC of 0.84, an accuracy of 80%, an F1 score of 0.67, a recall of 1.0 (Figure 6) and precision of 0.5. This indicates that half of the patients predicted by the model as responders were true responders, while all true responders in the cohort were selected given the recall of 1.0. A confusion matrix (Figure 7) was created to provide a visual representation of the model’s performance and to assess these classification metrics. To support clinical translation, the model was deployed as a website using Streamlit to allow users to enter data and estimate the probability of response to ICI treatment. 

Figure 6. ROC curves created using Matplotlib. Assesses the diagnostic performance of the true positive rate and false positive rate at all thresholds.

Figure 7. Confusion matrix created using Seaborn. Displays the prediction summary of the model for the testing set.

Discussion

Our study demonstrates how machine learning can be applied to improve patient selection for immune checkpoint inhibitor (ICI) therapy in non-small lung cancer (NSCLC). By using baseline clinical and biomarker data such as PD-L1 expression, TMB, and smoking history, our logistic regression model was able to differentiate between responders and non-responders with clinically meaningful precision. Specifically, the model showed strong performance through an ROC-AUC of 0.84, an F1 score of 0.67, and a recall of 1.0 on the test set. These findings suggest that machine learning can play a vital role in helping oncologists identify patients more likely to benefit from immunotherapy.

In traditional clinical settings, immune checkpoint inhibitor (ICI) therapy is administered to unselected cohorts, where only 20–30%1 of patients typically achieve an objective response based on RECIST criteria.18 Under our proposed model-guided selection, patients would first be filtered, and only those predicted as positive would undergo treatment. With our model achieving a precision of 0.50 and recall of 1.0, the expected proportion of true responders among treated patients would rise to 50% while identifying all true responders within the test dataset. While there are limitations to false positives inducing unnecessary toxicity and costs for non-responders, recall was favoured over precision as it suggests the model is effective at capturing all patients who are likely to respond to ICI therapy. Immunotherapy is often a secondary option to chemotherapy, meaning that falsely excluding potential responders from a potential live-saving therapy can be profoundly disadvantageous. Under model-guided selection, unlikely responders would first be filtered, and only those classed as positive would undergo treatment; based on our precision of 0.5, it is expected 50% of patients that underwent treatment would be true responders. In oncology, this research could translate into more efficient resource allocation, more effective treatment, and better outcomes for patients. Nevertheless, although the model was optimized using ROC-AUC, it can be tuned to favor different performance metrics, such as precision or F1 score, depending on the user’s intentions.

Our findings also confirm the relevance of many established and emerging features in predicting immunotherapy response. PD-L1 expression and TMB, widely known biomarkers, were influential predictors in our model, which aligns with previous research that shows incorporating these variables into screening protocols can improve patient selection and optimize ICI treatment strategies. However, these biomarkers have some limitations. PD-L1 suffers from spatial and temporal heterogeneity in its expression across tumour sites and TMB is unable to consistently predict responses across all cancer types6,7. In our model, we addressed these limitations by including additional factors such as neutrophil-to-lymphocyte ratio (NLR), pack-year smoking history, and microsatellite instability (MSI) to increase predictive power. These variables may capture additional components of the patient's tumour microenvironment and immune status, enhancing the model's performance.

Interestingly, the use of pack-years as a biomarker emerged as one of the most predictive features. This may reflect the biological mechanisms in which smoke-induced mutagenesis increases tumour antigenicity, making patients more responsive to immune-based therapies. This supports prior evidence of smoking status increasing TMB, leading to a higher response rate of PD-L1 inhibitors in NSCLC25,26. Furthermore, the use of NLR, a marker of systemic inflammation, was a biologically reasonable predictor27,28. Elevated NLR is associated with poor patient outcomes in various cancers, possibly due to immunosuppressive neutrophils impairing T-cell function.

Compared to more complex machine learning approaches, such as deep learning models used by Rakaee et al. or multimodal feature fusion used by Vanguri et al., our model may seem conservative12,13. However, logistic regression models offer transparency and reliability of data, which is important in clinical settings. Clinicians can trust models that provide clear coefficients, confidence intervals and variable importance. This is critical in high-stakes decision-making where black-box predictions may not be accepted without thorough review and validation.

Limitations

This study has notable limitations that should be considered alongside its findings. First, the small sample size (n = 224) is sufficient for logistic regression but may limit generalizability29. Additionally, the dataset was taken from a single institution (Memorial Sloan Kettering), which may introduce biases related to patient demographics, treatment plans, and data collection procedures. External validation through independent cohorts from multiple institutions will be essential to confirm the robustness and applicability of our model.

Another limitation is the use of purely clinical and molecular features. These features are low-cost and accessible compared to genomic sequencing or imaging, but lack complexity, which is needed to thoroughly analyze tumour microenvironments. Integrating additional data features such as RNA-sequencing (transcriptomics), proteomics, radiomics, and histopathology could significantly improve the predictive strength of the model30.

Future Directions

There are many future directions to be considered in regards to this study. Multiple studies have explored the use of multi-omics in oncology, and we plan to explore this further in the future31. Furthermore, an important consideration is model calibration. While our model performs well in terms of discrimination (ROC-AUC), future work should evaluate how its predicted probabilities align with the observed outcomes. Poorly calibrated models can over or underestimate clinical response. Techniques such as isotonic regression or Platt scaling may be employed to adjust probability outputs32,33.

Regarding evaluation metrics, ROC-AUC measures overall classification performance without considering class imbalance. Precision-recall AUC, calibration curves, and decision curve analysis could offer better insights34. Considering that there is a large class imbalance with only 30% responders, PR-AUC could offer informative insights as it focuses on the model's ability to correctly identify the minority class34. Furthermore, decision curve analysis can help evaluate the clinical benefit of model-guided patient selection, showing the real-life applications of this model.

Finally, the use of this model in Streamlit-based web applications demonstrates its real-world utility. While the app is currently only used for demonstration, it is the right step towards integration into clinical decision support systems (CDSS), where oncologists can enter baseline clinical data and receive response scores that can guide their treatment.

Conclusion

Ultimately, this study presents a clinically interpretable machine learning model designed to support patient selection and predict response to ICI therapy in NSCLC patients. By leveraging standard biomarkers like PD-L1 and TMB, our model demonstrated robust predictive accuracy. Notably, it achieved perfect sensitivity while delivering precision rates that substantially outperform historical cohort averages. These results demonstrate the potential of leveraging machine learning to improve patient selection for immunotherapy, enabling more personalized and efficient cancer treatment.

By using a logistic regression machine learning model, our approach is more effective for clinical setting compared to complex black box models, which lack interoperability. Furthermore, deploying the model as a web-based tool highlights its potential for integration into clinical decision support systems (CDSS). While the study is currently limited by data from a single institution and a modest sample size, it establishes a valuable foundation for future expansion into multi-omics data and external validation across diverse patient populations.

In conclusion, with further refinement and broader validation, this model could help oncologists identify patients who most likely could benefit from ICI therapy, ultimately contributing to more precise and equitable healthcare.




Conflicts of Interest Disclosure
There are no conflicts of interest to declare.


Web Resources
The program for this research can be accessed at https://immunotherapy-predictor.streamlit.app.


Data Availability
The data used for this study can be accessed through the cBioPortal which is publicly available at https://datacatalog.mskcc.org/dataset/10608.



References

  1. Ma, W., Xue, R., Zhu, Z. et al. Increasing cure rates of solid tumors by immune checkpoint inhibitors. Exp Hematol Oncol 12, 10 (2023). https://doi.org/10.1186/s40164-023-00372-8

  2. Barcellini, L., Nardin, S., Sacco, G. et al. Immune Checkpoint Inhibitors and Targeted Therapies in Early-Stage Non-Small-Cell Lung Cancer: State-of-the-Art and Future Perspectives. Cancers 17, 652 (2025). https://doi.org/10.3390/cancers17040652

  3. Ichihara, E., Harada, D., Inoue, K. et al. The impact of body mass index on the efficacy of anti-PD-1/PD-L1 antibodies in patients with non-small cell lung cancer. Lung Cancer 139, 140-145 (2020). https://doi.org/10.1016/j.lungcan.2019.11.011

  4. Tan, W. W. Non-small cell lung cancer (NSCLC). Medscape Updated March 31, 2026. https://emedicine.medscape.com/article/279960-overview

  5. Chen, Y. C., Chen, AY., Hong, R. et al. Optimizing T cell inflamed signature through a combination biomarker approach for predicting immunotherapy response in NSCLC. Sci Rep 14, 23617 (2024). https://doi.org/10.1038/s41598-024-82903-9

  6. Catalano, M., Iannone, L. F., Nesi, G. et al. Immunotherapy-related biomarkers: Confirmations and uncertainties. Crit Rev Oncol Hematol 192, 104135 (2023). https://doi.org/10.1016/j.critrevonc.2023.104135

  7. Bravaccini, S., Bronte, G. & Ulivi, P. TMB in NSCLC: A Broken Dream? Int J Mol Sci 22, 6536 (2021). https://doi.org/10.3390/ijms22126536

  8. Liao, J., Li, X., Gan, Y. et al. Artificial intelligence assists precision medicine in cancer treatment. Front Oncol 12, 998222 (2023). https://doi.org/10.3389/fonc.2022.998222

  9. Rezayi, S., Niakan Kalhori, S. R. & Saeedi, S. Effectiveness of Artificial Intelligence for Personalized Medicine in Neoplasms: A Systematic Review. Biomed Res Int 2022, 7842566 (2022). https://doi.org/10.1155/2022/7842566

  10. University of Oxford. Study demonstrates how AI can develop more personalised cancer treatment strategies. University of Oxford News June 17, 2024. https://www.ox.ac.uk/news/2024-06-17-study-demonstrates-how-ai-can-develop-more-personalised-cancer-treatment-strategies

  11. National Institutes of Health. AI tool predicts response to cancer therapy. NIH Research Matters January 22, 2025. https://www.nih.gov/news-events/nih-research-matters/ai-tool-predicts-response-cancer-therapy

  12. Rakaee, M., Adib, E., Ricciuti, B. et al. Association of Machine Learning-Based Assessment of Tumor-Infiltrating Lymphocytes on Standard Histologic Images With Outcomes of Immunotherapy in Patients With NSCLC. JAMA Oncol 9, 51-60 (2023). https://doi.org/10.1001/jamaoncol.2022.4933

  13. Vanguri, R. S., Luo, J., Aukerman, A. T. et al. Multimodal integration of radiology, pathology and genomics for prediction of response to PD-(L)1 blockade in patients with non-small cell lung cancer. Nat Cancer 3, 1151-1164 (2022). https://doi.org/10.1038/s41591-022-01976-9

  14. Hua, Y., Stead, T. S., George, A. et al. Clinical Risk Prediction with Logistic Regression: Best Practices, Validation Techniques, and Applications in Medical Research. Acad Med Surg 1, 131964 (2025). https://doi.org/10.62186/001c.131964

  15. Memorial Sloan Kettering Cancer Center. Non-Small Cell Lung Cancer (MSKCC, J Clin Oncol 2018): IMPACT sequencing of 240 NSCLC tumor/normal pairs treated at MSKCC with anti-PD-(L)1 based therapy. MSK Data Catalog 10608 (2018). https://datacatalog.mskcc.org/dataset/10608

  16. Wang, J., Zeng, Z., Li, Z. et al. The clinical application of artificial intelligence in cancer precision treatment. J Transl Med 23, 120 (2025). https://doi.org/10.1186/s12967-025-06139-5

  17. Rizvi, H., Sanchez-Vega, F., La, K. et al. Molecular Determinants of Response to Anti-Programmed Cell Death-1 and Anti-Programmed Death-Ligand 1 Blockade in Patients with Non-Small-Cell Lung Cancer Profiled by Targeted Next-Generation Sequencing. J Clin Oncol 36, 633-641 (2018). https://doi.org/10.1200/JCO.2017.75.3384

  18. Eisenhauer, E. A., Therasse, P., Bogaerts, J. et al. New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1). Eur J Cancer 45, 228–247 (2009). https://doi.org/10.1016/j.ejca.2008.10.026.

  19. Chawla, N. V., Bowyer, K. W., Hall, L. O. et al. SMOTE: synthetic minority over-sampling technique. J Artif Intell Res16, 321–357 (2002). https://doi.org/10.1613/jair.953

  20. Python Software Foundation. Python programming language, version 3.9.6. 2021. Available from: python.org.

  21. The Pandas Development Team. pandas-dev/pandas: Pandas. Zenodo; 2020. https://doi.org/10.5281/zenodo.3509134

  22. Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825–2830.

  23. Waskom ML. seaborn: statistical data visualization. J Open Source Softw. 2021;6(60):3021. https://doi.org/10.21105/joss.03021

  24. Hunter JD. Matplotlib: a 2D graphics environment. Comput Sci Eng. 2007;9(3):90–95. https://doi.org/10.1109/MCSE.2007.55

  25. Wang, X., Ricciuti, B., Nguyen, T. et al. Association between Smoking History and Tumor Mutation Burden in Advanced Non-Small Cell Lung Cancer. Cancer Res 81, 2566-2573 (2021). https://doi.org/10.1158/0008-5472.CAN-20-3991

  26. Sun, L. Y., Cen, W. J., Tang, W. T. et al. Smoking status combined with tumor mutational burden as a prognosis predictor for combination immune checkpoint inhibitor therapy in non-small cell lung cancer. Cancer Med 10, 6610-6617 (2021). https://doi.org/10.1002/cam4.4197

  27. Mosca, M., Nigro, M. C., Pagani, R. et al. Neutrophil-to-Lymphocyte Ratio (NLR) in NSCLC, Gastrointestinal, and Other Solid Tumors: Immunotherapy and Beyond. Biomolecules 13, 1803 (2023). https://doi.org/10.3390/biom13121803

  28. Valero, C., Lee, M., Hoen, D. et al. Pretreatment neutrophil-to-lymphocyte ratio and mutational burden as biomarkers of tumor response to immune checkpoint inhibitors. Nat Commun 12, 729 (2021). https://doi.org/10.1038/s41467-021-20935-9

  29. Rajput, D., Wang, W. J. & Chen, C. C. Evaluation of a decided sample size in machine learning applications. BMC Bioinformatics 24, 48 (2023). https://doi.org/10.1186/s12859-023-05156-9

  30. Hasin, Y., Seldin, M. & Lusis, A. Multi-omics approaches to disease. Genome Biol 18, 83 (2017). https://doi.org/10.1186/s13059-017-1215-1.

  31. Raufaste-Cazavieille, V., Santiago, R. & Droit, A. Multi-omics analysis: Paving the path toward achieving precision medicine in cancer treatment and immuno-oncology. Front Mol Biosci 9, 962743 (2022). https://doi.org/10.3389/fmolb.2022.962743

  32. Van Calster, B., McLernon, D. J., van Smeden, M. et al. Calibration: the Achilles heel of predictive analytics. BMC Med17, 230 (2019). https://doi.org/10.1186/s12916-019-1466-7

  33. Van Houwelingen, J. C. & Le Cessie, S. Predictive value of statistical models. Stat Med 9, 1303-1325 (1990). https://doi.org/10.1002/sim.4780091109

  34. Brownlee, J. How to Calibrate Probabilities for Imbalanced Classification. Machine Learning Mastery August 21, 2020. https://machinelearningmastery.com/how-to-calibrate-probabilities-for-imbalanced-classification/