Book Details

A COMPARATIVE STUDY OF MACHINE LEARNING ALGORITHMS FOR CUSTOMER CHURN PREDICTION

International Journal of Computer Science (IJCS) Published by SK Research Group of Companies (SKRGC)

Download this PDF format

Abstract

Customer churn, the loss of subscribers to a competing service, is one of the most pressing concerns for subscription-driven businesses such as telecommunication operators, banks, and streaming platforms. Because retaining an existing customer is substantially cheaper than acquiring a new one, accurately identifying customers who are likely to churn allows firms to intervene with targeted retention offers before the relationship ends. This paper presents an extended comparative study of eleven supervised machine learning algorithms—Logistic Regression, Decision Tree, Random Forest, Extra Trees, Gradient Boosting, AdaBoost, XGBoost, Support Vector Machine, k-Nearest Neighbors, Naive Bayes, and a Multi-Layer Perceptron neural network—applied to the task of customer churn prediction on a representative telecommunications dataset. Beyond a single train/test comparison, the study evaluates each model using stratified 5-fold cross-validation, paired statistical significance testing, hyperparameter tuning via grid search, sensitivity to class-imbalance handling, learning-curve analysis, and computational-complexity considerations. Each model is scored on accuracy, precision, recall, F1-score, and area under the ROC curve (AUC). Experimental results show that boosted ensembles—Gradient Boosting and AdaBoost—achieve the strongest and statistically indistinguishable top-tier performance, significantly outperforming Random Forest, SVM, Naive Bayes, k-NN, the neural network, and the Decision Tree under cross-validation (paired t-test, p < 0.05). Feature-importance analysis further reveals that contract type, tenure, and monthly charges are the strongest predictors of churn. These findings offer practical, statistically grounded guidance for businesses seeking to deploy interpretable and effective churn-prediction pipelines.

References

  1. L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
  2. J. H. Friedman, "Greedy function approximation: A gradient boosting machine," Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001.
  3. C. Cortes and V. Vapnik, "Support-vector networks," Machine Learning, vol. 20, no. 3, pp. 273–297, 1995.
  4. T. Cover and P. Hart, "Nearest neighbor pattern classification," IEEE Trans. Inf. Theory, vol. 13, no. 1, pp. 21–27, 1967.
  5. T. Vafeiadis, K. I. Diamantaras, G. Sarigiannidis, and K. Ch. Chatzisavvas, "A comparison of machine learning techniques for customer churn prediction," Simulation Modelling Practice and Theory, vol. 55, pp. 1–9, 2015.
  6. A. Idris, A. Khan, and Y. S. Lee, "Genetic programming and adaboosting based churn prediction for telecom," in Proc. IEEE Int. Conf. Systems, Man, and Cybernetics (SMC), 2012.
  7. A.K. Ahmad, A. Jafar, and K. Aljoumaa, "Customer churn prediction in telecom using machine learning in big data platform," Journal of Big Data, vol. 6, no. 1, 2019.
  8. W. Verbeke, K. Dejaeger, D. Martens, J. Hur, and B. Baesens, "New insights into churn prediction in the telecommunication sector: A profit driven data mining approach," European Journal of Operational Research, vol. 218, no. 1, pp. 211–229, 2012.
  9. D. E. Rumelhart, G. E. Hinton, and R. J. Williams, "Learning representations by back-propagating errors," Nature, vol. 323, pp. 533–536, 1986.
  10. I. Rish, "An empirical study of the naive Bayes classifier," in IJCAI Workshop on Empirical Methods in AI, 2001.
  11. F. Pedregosa et al., "Scikit-learn: Machine learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
  12. N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002.
  13. T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proc. ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785–794.
  14. Y. Freund and R. E. Schapire, "A decision-theoretic generalization of on-line learning and an application to boosting," Journal of Computer and System Sciences, vol. 55, no. 1, pp. 119–139, 1997.
  15. P. Geurts, D. Ernst, and L. Wehenkel, "Extremely randomized trees," Machine Learning, vol. 63, no. 1, pp. 3–42, 2006.

Keywords

Customer churn, machine learning, classification, predictive analytics, random forest, gradient boosting, telecommunications.

Image
  • Format Volume 14, Issue 2, No 01, 2026
  • Copyright All Rights Reserved ©2026
  • Year of Publication 2026
  • Author R. Muneeswari, C Jebatheeswari
  • Reference IJCS-728
  • Page No 056-070

Copyright 2026 SK Research Group of Companies. All Rights Reserved.