Development of a Data Mining-Based Risk Screening Model for Non-Communicable Diseases among Older Adults to Support Community Health Surveillance

Authors

  • Anucha Ruangsiriwattanakul Faculty of Science and Technology, Uttaradit Rajabhat University
  • Thitiporn Tienkanitikul Faculty of Science and Technology, Uttaradit Rajabhat University
  • Tanakorn Tanawat Faculty of Science and Technology, Uttaradit Rajabhat University
  • Chanida Ruangsiriwattanakul Faculty of Science and Technology, Uttaradit Rajabhat University

DOI:

https://doi.org/10.57260/stc.2026.1456

Keywords:

Non-communicable disease risk screening, Diabetes mellitus, Hypertension, Older adults, Machine learning

Abstract

In 2025, Thailand entered an Aged Society, with adults aged 60 years and older comprising over 20% of the total population, resulting in a growing burden of non-communicable diseases (NCDs), particularly diabetes mellitus and hypertension, which have become major public health concerns, especially in community settings with limited resources and data-driven decision-support tools. This study aimed to develop a risk screening model for non-communicable diseases among older adults using data mining and machine learning techniques applied to community health data. The study analyzed health screening records from 3,419 older adults in Laplae District, Uttaradit Province, Thailand, collected between 2022 and 2025. Data were processed following the CRISP-DM framework. Random Forest and Extreme Gradient Boosting (XGBoost) models were developed and evaluated using Accuracy, Precision, Recall, F1-score, and Confusion Matrix.
The results demonstrated that XGBoost was more appropriate for public health screening purposes, achieving a Sensitivity of 25.0% for diabetes risk detection and 16.0% for hypertension, compared to 17.3% and 4.0% for Random Forest, respectively. Although Random Forest yielded higher overall accuracy (91.3% and 91.7%), XGBoost produced fewer false negatives, which is critically important in preventive health screening contexts. Key risk-associated factors included body mass index, waist circumference, blood glucose level, blood pressure, and age. The findings suggest that integrating machine learning with community-level health data can effectively support preliminary risk screening, proactive health surveillance, and local healthcare planning for older adults. The model is intended as a decision-support tool only, and is not a substitute for clinical diagnosis.

Downloads

Download data is not yet available.

References

กรมควบคุมโรค. (2566). รายงานสถานการณ์โรคไม่ติดต่อเรื้อรังของประเทศไทย. กระทรวงสาธารณสุข.

สำนักงานสถิติแห่งชาติ. (2567). รายงานสถานการณ์ผู้สูงอายุในประเทศไทย. สำนักงานสถิติแห่งชาติ.

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the

nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.

(pp. 785–794). https://doi.org/10.1145/2939672.2939785

Deberneh, H. M., & Kim, I. (2021). Prediction of type 2 diabetes based on machine learning algorithm. International Journal of Environmental Research and Public Health, 18(6),

https://doi.org/10.3390/ijerph18063317

Han, J., Kamber, M., & Pei, J. (2011). Data mining: Concepts and techniques. (3rd ed.). Morgan Kaufmann.

Hu, B., Wang, C., Jiang, K., Shen, Z., Yang, X., Yin, M., Liang, B., Xie, Q., Ye, Y., & Gao, Z. (2021). Development and validation of a novel diagnostic model for initially clinical diagnosed gastrointestinal stromal tumors using an extreme gradient-boosting machine. BMC Gastroenterol, 21, 481. https://doi.org/10.1186/s12876-021-02048-1

Islam, S. M. S., Talukder, A., Awal, M. A., Siddiqui, M. M. U., Ahamad, M. M., Ahammed, B., Rawal, L. B., Alizadehsani, R., Abawajy, J., Laranjo, L., Chow, C. K., & Maddison, R. (2022). Machine learning approaches for predicting hypertension and its associated factors using population-level data from three South Asian countries. Frontiers in Cardiovascular Medicine, 9, 839379. https://doi.org/10.3389/fcvm.2022.839379

Powers, D. M. W. (2011). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. Journal of Machine Learning Technologies, 2(1), 37–63. https://doi.org/10.48550/arXiv.2010.16061

Wang, L., Wang, X., Chen, A., Jin, X., & Che, H. (2020). Prediction of type 2 diabetes risk and its effect evaluation based on the XGBoost model. Healthcare, 8(3), 247. https://doi.org/10.3390/healthcare8030247

Wirth, R., & Hipp, J. (2000). CRISP-DM: Towards a standard process model for data mining. Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining. (pp. 29–39). https://www.cs.unibo.it/~danilo.montesi/CBD/Beatriz/10.1.1.198.5133.pdf

World Health Organization. (2014). Global status report on noncommunicable diseases 2014. https://www.who.int/publications/i/item/9789241564854

Downloads

Additional Files

Published

2026-08-31

How to Cite

Ruangsiriwattanakul, A., Tienkanitikul, T., Tanawat , T., & Ruangsiriwattanakul, C. (2026). Development of a Data Mining-Based Risk Screening Model for Non-Communicable Diseases among Older Adults to Support Community Health Surveillance. Science and Technology to Community, 4(4), 32–46. https://doi.org/10.57260/stc.2026.1456

Issue

Section

Research Articles