Customer Churn Prediction Using Machine Learning Models: A Comparative Study Using R Software

Main Article Content

Uppu Venkata Subbarao
Tedlapu Narayana Rao
Vantaku Bala
K.B Rajeswara Rao

Abstract

Predicting customer churn is essential for improving retention and supporting long-term business growth. In this study, we compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R. Our approach included data preprocessing, exploratory analysis, model development, performance evaluation, and further analysis. We developed and evaluated four classification algorithms: Logistic Regression, Decision Tree, Random Forest, and Extreme Gradient Boosting, using an 80:20 train-test split. We assessed each model's accuracy, precision, recall, and F1-score. Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset. Feature importance analysis indicated that contract type, customer tenure, monthly charges, total charges, and internet service were the key factors influencing churn. These results suggest that explainable machine learning offers both strong predictive performance and greater transparency. The R-based framework we present provides a practical, reproducible approach to support customer retention strategies and help managers make evidence-based decisions in customer relationship management.

Downloads

Download data is not yet available.

Article Details

Section

Articles

How to Cite

[1]
Uppu Venkata Subbarao, Tedlapu Narayana Rao, Vantaku Bala, and K.B Rajeswara Rao , Trans., “Customer Churn Prediction Using Machine Learning Models: A Comparative Study Using R Software”, IJMH, vol. 12, no. 12, pp. 5–13, Aug. 2026, doi: 10.35940/ijmh.L1893.12120826.
Share |

References

Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160. DOI: 10.1109/ACCESS.2018.2870052

Alloghani, M., Al-Jumeily, D., Hussain, A., Aljaaf, A. J., Mustafina, J., & Petrov, E. (2020). Decision tree algorithms: A survey. In D. Al-Jumeily et al. (Eds.), Artificial Intelligence Review. Springer. https://www.academia.edu/97048115/

Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. DOI:10.1016/j.inffus.2019.12.012

Bzdok, D., Altman, N., & Krzywinski, M. (2018). Statistics versus machine learning. Nature Methods, 15(4), 233–234.DOI:10.1038/nmeth.4642

Carvalho, D. V., Pereira, E. M., & Cardoso, J. S. (2019). Machine learning interpretability: A survey on methods and metrics. Electronics, 8(8), Article 832. DOI: 10.3390/electronics8080832

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). DOI: 10.1145/2939672.2939785

Chen, T., He, T., Benesty, M., Khotilovich, V., Tang, Y., Cho, H., Mitchell, R., Cano, I., Zhou, T., Li, M., Xie, J., Lin, M., Geng, Y., & Li, Y. (2024). xgboost: Extreme Gradient Boosting (R package Version 1.7.10.1). https://CRAN.R-project.org/

Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21, Article 6. DOI: 10.1186/s12864-019-6413-7

Duan, Y., Edwards, J. S., & Dwivedi, Y. K. (2019). Artificial intelligence for decision making in the era of big data—Evolution, challenges and research agenda. International Journal of Information Management, 48, 63–71. DOI: 10.1016/j.ijinfomgt.2019.01.021

Géron, A. (2023). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly Media.

Hastie, T., Tibshirani, R., & Friedman, J. (2021). The elements of statistical learning: Data mining, inference, and prediction (2nd corrected ed.). Springer. DOI: 10.1007/978-0-387-84858-7

IBM. (2013). Telco customer churn [Data set]. Kaggle. https://www.kaggle.com. Work remains significant, see the declaration

James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: With applications in R (2nd ed.). Springer. DOI: 10.1007/978-1-0716-1418-1

Kotu, V., & Deshpande, B. (2019). Data science: Concepts and practice (2nd ed.). Morgan Kaufmann. DOI: 10.1016/C2018-0-01769-6

Kumar, V., & Reinartz, W. (2016). Creating enduring customer value. Journal of Marketing, 80(6), 36–68. 10.1509/jm.15.0414

Lemon, K. N., & Verhoef, P. C. (2016). Understanding customer experience throughout the customer journey. Journal of Marketing, 80(6), 69–96. DOI: 10.1509/jm.15.0420

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30. https://proceedings.neurips.cc/paper/2017/

Manzoor, A., Qureshi, M. A., Kidney, E., & Longo, L. (2024). A review on machine learning methods for customer churn prediction and recommendations for business practitioners. IEEE Access, 12, 70434–70463. DOI: 10.1109/ACCESS.2024.3402092

Molnar, C. (2022). Interpretable machine learning (2nd ed.). Leanpub. https://christophm.github.io/interpretable-ml-book/

Probst, P., Wright, M. N., & Boulesteix, A.-L. (2019). Hyperparameters and tuning strategies for random forest. WIREs Data Mining and Knowledge Discovery, 9(3), e1301. DOI: 10.1002/widm.1301

R Core Team. (2025). R: A language and environment for statistical computing (Version 4.5.1) [Computer software]. R Foundation for Statistical Computing. https://www.R-project.org/

Sarker, I. H. (2022). Machine learning: Algorithms, real-world applications and research directions. SN Computer Science, 3, Article 160. DOI: 10.1007/s42979-021-00865-8

Therneau, T., & Atkinson, B. (2023). rpart: Recursive partitioning and regression trees (R package Version 4.1-21). https://CRAN.R-project.org/package=rpart