Abstract

Gradient-boosting classifiers now underpin most screening tools proposed for postpartum mental health, but published comparisons rarely separate what the modelling pipeline contributes from what the clinical construct itself contributes. Here we benchmark XGBoost, CatBoost, and LightGBM inside a single leakage-controlled pipeline, asking whether performance gaps trace back to identifiable structural properties rather than inconsistent data handling. A public postpartum survey dataset of 1,503 records was split 80:20 into development and test sets, with encoding and imputation fitted on the development set alone, and each classifier tuned via stratified five-fold cross-validation over a shared search space, isolating architecture as the sole variable. LightGBM's histogram-based, leaf-wise growth produced the strongest held-out discrimination (accuracy 88.4%, weighted F1 88.4%, balanced accuracy 87.8%, macro F1 88.0%), ahead of XGBoost (86.0%) and CatBoost (82.4%), with its advantage concentrated in the Moderate and High categories (81 of 92 High-category records correctly classified, versus 80 and 76). The selected classifier was integrated into a four-layer architecture, preprocessing, predictive inference, dialogue management, and deterministic recommendation, keeping the model's output separate from user-facing text and independently auditable. We report the full search space and the structural mechanisms plausibly explaining the ranking, and state plainly that training time, memory, and inference latency were not measured, naming this as required future work under TRIPOD+AI. The results demonstrate a reproducible method for comparing boosting architectures on a fixed screening task, they do not establish diagnostic validity, calibrated suicide-risk estimation, or clinical readiness, which require external validation beyond this internal benchmark.

Keywords

Gradient Boosting, Xgboost, Lightgbm, Catboost, Hyperparameter Optimization, Postpartum Mental Health, Algorithmic Architecture, Computational Performance, Safety-Aware Artificial Intelligence, Conversational Agent,

Downloads

Download data is not yet available.

References

  1. Y. Yildiz, A. Kalayci, Gradient Boosting Decision Trees on Medical Diagnosis over Tabular Data. IEEE International Conference on AI and Data Analytics (ICAD), IEEE, USA. https://doi.org/10.1109/ICAD65464.2025.11114069
  2. R. Zhang, Y. Liu, Z. Zhang, R. Luo, B. Lv. Interpretable Machine Learning Model for Predicting Postpartum Depression: Retrospective Study. JMIR Medical Informatics, 13, (2025) e58649. https://doi.org/10.2196/58649
  3. W. Qi, Y. Wang, Y. Wang, S. Huang, C. Li, H. Jin, J. Zuo, X. Cui, Z. Wei, Q. Guo, J. Hu. Prediction of Postpartum Depression in Women: Development and Validation of Multiple Machine Learning Models. Journal of Translational Medicine, 23, (2025) 291. https://doi.org/10.1186/s12967-025-06289-6
  4. Z. Ma, M. Horvath, D.M. Stamilio, K. Sekyere, M.N. Gurcan. Building a Machine Learning Model to Predict Postpartum Depression from Electronic Health Records in a Tertiary Care Setting. Journal of Clinical Medicine, 14, (2025) 6644. https://doi.org/10.3390/jcm14186644
  5. X. Huang, L. Zhang, C. Zhang, J. Li, C. Li. Postpartum Depression Risk Prediction Using Explainable Machine Learning Algorithms. Frontiers of Medicine, 12, (2025) 1565374. http://dx.doi.org/10.3389/fmed.2025.1565374
  6. H. Kim. Predictive Analysis of Postpartum Depression Using Machine Learning. Healthcare, 13, (2025) 897. https://doi.org/10.3390/healthcare13080897
  7. C. Sona Ajay, S. Juliet, A. Alex. Machine Learning and Survival Analysis Models for Postpartum Depression: A Comprehensive Risk Factor Analysis. In Proceedings of the 2024 10th International Conference on Advanced Computing and Communication Systems (ICACCS), India. https://doi.org/10.1109/ICACCS60874.2024.10717207
  8. S. García-Méndez, F. de Arriba-Pérez. Detecting and Explaining Postpartum Depression in Real-Time with Generative Artificial Intelligence. Applied Artificial Intelligence, 39, (2025) 2515063. https://doi.org/10.1080/08839514.2025.2515063
  9. V.S.F. Gonçalves, V.R. de Carvalho. A Review of Interpretability Methods for Gradient Boosting Decision Trees. Journal of the Brazilian Computer Society, 31, (2025) 639–653. https://doi.org/10.5753/jbcs.2025.5324
  10. H. Semmelrock, T. Ross-Hellauer, S. Kopeinik, D. Theiler, A. Haberl, S. Thalmann, D. Kowald. Reproducibility in Machine-Learning-Based Research: Overview, Barriers, and Drivers. AI Magazine, 46(2), (2025) e70002. https://doi.org/10.1002/aaai.70002
  11. M. Alkhateeb, A. Nayeem, A. Ahmed, M. Alsahli, J. Sheikh, A. Abd-Alrazaq. AI for Detecting and Predicting Postpartum Depression: Scoping Review. Journal of Medical Internet Research, 28, (2026) e77376. https://doi.org/10.2196/77376
  12. U.S. Anaduaka, A.O. Oladosu, S. Katsande, C.S. Frempong, S. Awuku-Amador. Leveraging Artificial Intelligence in the Prediction, Diagnosis and Treatment of Depression and Anxiety among Perinatal Women in Low- and Middle-Income Countries: A Systematic Review. BMJ Mental Health, 28, (2025) e301445. https://doi.org/10.1136/bmjment-2024-301445
  13. J. Xia, C. Chen, X. Lu, T. Zhang, T. Wang, Q. Wang, Q. Zhou, Artificial Intelligence-Oriented Predictive Model for the Risk of Postpartum Depression: A Systematic Review. Front. Public Health, 13, (2025) 1631705. https://doi.org/10.3389/fpubh.2025.1631705
  14. L. Sibbald, M.I. van den Heuvel, M.R. Haas, C.J. van Lissa, H.J.A. van Bakel, J. Jongerling, L. P. Hulsbosch, L. Muskens, M. G. B. M. Boekhorst, I. Schwabe. Identifying Prenatal Risk Factors of Postpartum Depression with Machine Learning. Scientific Reports, 15, (2025) 34610. https://doi.org/10.1038/s41598-025-18204-6
  15. Y. Xie, H. Zheng, W. Gan, C. Su, M. Shams, J. Yang. The Performance of Machine Learning Models in Predicting Postpartum Depression: A Meta-Analysis and Systematic Review. Journal of Reproductive and Infant Psychology, 43(5), (2025) 1093–1110. https://doi.org/10.1080/02646838.2025.2517103
  16. L. Ou, Q. Shen, M. Xiao, W. Wang, T. He, B. Wang. Prevalence of Co-Morbid Anxiety and Depression in Pregnancy and Postpartum: A Systematic Review and Meta-Analysis. Psychological Medicine, 55, (2025) e93. https://doi.org/10.1017/S0033291725000601
  17. K. Zivin, C. Zhong, A. Rodríguez-Putnam, E. Spring, Q. Cai, A. Miller, L. J. Johns, V. A. Kalesnikava, A. Courant, B. Mezuk. Suicide Mortality during the Perinatal Period. JAMA Netw Open, 7, (2024) e2418887. https://doi.org/10.1001/jamanetworkopen.2024.18887
  18. H. Yu, Q. Shen, E. Bränn, Y. Yang, A.S. Oberg, U.A. Valdimarsdóttir, D. Lu. Perinatal Depression and Risk of Suicidal Behavior. JAMA Netw. Open, 7, (2024) e2350897. https://doi.org/10.1001/jamanetworkopen.2023.50897
  19. S. Goldman-Mellor, M. Olfson, A. Gemmill, C. Margerison. Incidence and Risk Factors for Suicide Attempt during Pregnancy and the Postpartum Period. The Journal of Clinical Psychiatry, 86, (2025) 24m15633. https://doi.org/10.4088/JCP.24m15633
  20. J. M. Martínez-Ramírez, R.A. Peinado-Molina, A. Hernández-Martínez, J.M. Martínez-Galiano. Predictive Models of Suicidal Ideation Risk in Perinatal-Stage Women Based on Sociodemographic and Clinical Data. Frontiers in Artificial Intelligence Medicine and Public Health, 9, (2026) 1774453. https://doi.org/10.3389/frai.2026.1774453
  21. Y. Hua, S. Siddals, Z. Ma, I. Galatzer-Levy, W. Xia, C. Hau, H. Na, M. Flathers, J. Linardon, C. Ayubcha, J. Torous. Charting the Evolution of Artificial Intelligence Mental Health Chatbots from Rule-Based Systems to Large Language Models: A Systematic Review. World Psychiatry, 24(3), (2025) 383–394. https://doi.org/10.1002/wps.21352
  22. A. Arnaiz-Rodriguez, M. Baidal, E. Derner, J. Layton Annable, M. Ball, M. Ince, E. Perez Vallejos, N. Oliver, Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models. JMIR Mental Health, 13(1), (2026) e88435. https://doi.org/10.2196/88435
  23. A.-R. Wrightson-Hester, G. Anderson, J. Dunstan, P. M. McEvoy, C.J. Sutton, B. Myers, S.J. Egan, S.J. Tai, M. Johnston-Hollitt, W. Chen, T. Gedeon, J.C. Moullin, W. Mansell, A Rule-Based Conversational Agent for Mental Health and Well-Being in Young People: Formative Case Series during the Rise of Generative AI. JMIR Formative Research, 9, (2025) e69841. https://doi.org/10.2196/69841
  24. G.S. Collins, K.G.M. Moons, P. Dhiman, R.D. Riley, A.L. Beam, B. van Calster, M. Ghassemi, X. Liu, J.B. Reitsma, M. van Smeden, A.L. Boulesteix, TRIPOD+AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods. BMJ, 385, (2024) e078378. https://doi.org/10.1136/bmj-2023-078378
  25. C. Meaney, X. Wang, J. Guan, T.A. Stukel, Comparison of Methods for Tuning Machine Learning Model Hyper-Parameters: with Application to Predicting High-Need High-Cost Health Care Users. BMC Medical Research Methodology, 25(1), (2025) 134. https://doi.org/10.1186/s12874-025-02561-x
  26. V. Gupta, S. Tripathi, D. Singh, A. Bansal, (2024) Predictive Algorithms for Early Postpartum Depression Detection: CatBoost vs. LightGBM. In Proceedings of the 2024 11th International Conference on Reliability, Infocom Technologies and Optimization (ICRITO), IEEE, Noida, India. https://doi.org/10.1109/ICRITO61523.2024.10522300
  27. N. Hagatulah, E. Bränn, A.S. Oberg, U.A. Valdimarsdóttir, D. Lu, Q. Shen, Perinatal Depression and Risk of Mortality: Nationwide, Register Based Study in Sweden. BMJ, 384, (2024) e075462. https://doi.org/10.1136/bmj-2023-075462
  28. Z. Kaminsky, R. J. McQuaid, K. G. Hellemans, Z. R. Patterson, M. Saad, R. L. Gabrys, T. Kendzerska, A. Abizaid, R. Robillard, Machine Learning-Based Suicide Risk Prediction Model for Suicidal Trajectory on Social Media Following Suicidal Mentions: Independent Algorithm Validation. Journal of Medical Internet Research, 26, (2024) e49927. https://doi.org/10.2196/49927
  29. L. Liu, Z. Li, Y. Hu, C. Li, S. He, S. Zhang, J. Gao, H. Zhu, G. Huang, Predictive Performance of Machine Learning for Suicide in Adolescents: Systematic Review and Meta-Analysis. Journal of Medical Internet Research, 27, (2025) e73052. https://doi.org/10.2196/73052
  30. R. Li, Y. Yue, X. Gu, L. Xiong, M. Luo, L. Li, Risk Prediction Models for Adolescent Suicide: A Systematic Review and Meta-Analysis. Psychiatry Research, 347, (2025) 116405. https://doi.org/10.1016/j.psychres.2025.116405
  31. L. Deng, K. Lu, H. Hu, An Interpretable LightGBM Model for Predicting Coronary Heart Disease: Enhancing Clinical Decision-Making with Machine Learning. PLoS one, 20(9), (2025) e0330377. https://doi.org/10.1371/journal.pone.0330377
  32. J. Xu, T. Chen, X. Fang, L. Xia, X. Pan, Prediction Model of Pressure Injury Occurrence in Diabetic Patients during ICU Hospitalization: XGBoost Machine Learning Model Can Be Interpreted Based on SHAP. Intensive and Critical Care Nursing, 83, (2024) 103715. https://doi.org/10.1016/j.iccn.2024.103715
  33. M.R. Khurshid, S. Manzoor, T. Sadiq, L. Hussain, M.S. Khan, A.K. Dutta, Unveiling Diabetes Onset: Optimized XGBoost with Bayesian Optimization for Enhanced Prediction. PLoS ONE, 20, (2025) e0310218. https://doi.org/10.1371/journal.pone.0310218
  34. C. Yang, E.A. Fridgeirsson, J.A. Kors, J.M. Reps, P.R. Rijnbeek, Impact of Random Oversampling and Random Undersampling on the Performance of Prediction Models Developed Using Observational Health Data. Journal of Big Data, 11(1), (2024) 7. https://doi.org/10.1186/s40537-023-00857-7
  35. J. Digitale, D. Franzon, M.J. Pletcher, C.E. McCulloch, E.D. Gennatas, Methods for Addressing Missingness in Electronic Health Record Data for Clinical Prediction Models: Comparative Evaluation. JMIR Medical Informatics, 13(1), (2025) e79307. https://doi.org/10.2196/79307
  36. R. Sanjeewa, R. Iyer, P. Apputhurai, N. Wickramasinghe, D. Meyer, Empathic Conversational Agent Platform Designs and Their Evaluation in the Context of Mental Health: Systematic Review. JMIR Mental Health, 11, (2024) e58974. https://doi.org/10.2196/58974
  37. W. Pichowicz, M. Kotas, P. Piotrowski, Performance of Mental Health Chatbot Agents in Detecting and Managing Suicidal Ideation. Scientific Reports, 15(1), (2025) 31652. https://doi.org/10.1038/s41598-025-17242-4
  38. A. Palmer, D. Schwan, Digital Mental Health Tools and AI Therapy Chatbots: A Balanced Approach to Regulation. Hastings Center Report, 55(3), (2025) 15–29. https://doi.org/10.1002/hast.4979
  39. Y. Jin, J. Liu, P. Li, B. Wang, Y. Yan, H. Zhang, C. Ni, J. Wang, Y. Li, Y. Bu, Y. Wang, The Applications of Large Language Models in Mental Health: Scoping Review. Journal of Medical Internet Research, 27, (2025) e69284. https://doi.org/10.2196/69284
  40. Y. Hua, H. Na, Z. Li, et al., A Scoping Review of Large Language Models for Generative Tasks in Mental Health Care. npj digital medicine, 8, (2025) 230. https://doi.org/10.1038/s41746-025-01611-4
  41. Stolltho, Postpartum Depression Classification [Kaggle Notebook and Associated Input Data]. Kaggle, n.d. Available online: https://www.kaggle.com/code/stolltho/postpartum-depression-classification/input
  42. E. Zain, Y. Watanabe, S. Takabayashi, L. Por, S. Fujita, S. Moriyama, A. Honma, N. Fukui, S. Boku, Psychometric Evaluation of the Japanese Edinburgh Postnatal Depression Scale for Screening Postpartum Anxiety. Frontiers in Psychiatry, 16, (2025) 1659497. https://doi.org/10.3389/fpsyt.2025.1659497
  43. L. Liou, E. Scott, P. Parchure, Y. Ouyang, N. Egorova, R. Freeman, I. S. Hofer, G. N. Nadkarni, P. Timsina, A. Kia, M. A. Levin, Assessing Calibration and Bias of a Deployed Machine Learning Malnutrition Prediction Model within a Large Healthcare System. npj Digit. Med., 7, (2024) 141. https://doi.org/10.1038/s41746-024-01141-5
  44. I. Partheniadis, P. Talimtzi, A. Nikolakopoulou, A.-B. Haidich, Machine Learning-Based COVID-19 Prognostic Models Lag Behind in Reporting Quality: Findings from a TRIPOD/TRIPOD+AI Systematic Review. Diagn. Progn. Res., 10, (2026) 3. https://doi.org/10.1186/s41512-026-00218-x
  45. A. J. Millner, M. D. Lee, M. K. Nock, Single-Item Measurement of Suicidal Behaviors: Validity and Consequences of Misclassification. PLoS ONE, 10, (2015) e0141606. https://doi.org/10.1371/journal.pone.0141606
  46. S.E. Davis, C. Dorn, D.J. Park, M.E. Matheny, Emerging Algorithmic Bias: Fairness Drift as the Next Dimension of Model Maintenance and Sustainability. Journal of the American Medical Informatics Association, 32, (2025) 845–854. https://doi.org/10.1093/jamia/ocaf039
  47. K.G.M. Moons, J.A.A. Damen, T. Kaul, L. Hooft, C. Andaur Navarro, P. Dhiman, A.L. Beam, B. Van Calster, L.A. Celi, S. Denaxas, A.K. Denniston, M. Ghassemi, G. Heinze, A.P. Kengne, L. Maier-Hein, X. Liu, P. Logullo, M.D McCradden, N. Liu, L. Oakden-Rayner, ,K. Singh, D.S Ting, ,L. Wynants, ,B. Yang, ,J.B Reitsma, ,R.D Riley, G.S Collins, M.V. Smeden, PROBAST+AI: An Updated Quality, Risk of Bias, and Applicability Assessment Tool for Prediction Models Using Regression or Artificial Intelligence Methods. BMJ, 388, (2025) e082505. https://doi.org/10.1136/bmj-2024-082505
  48. Q. Abbas, W. Jeong, S.W. Lee, Explainable AI in Clinical Decision Support Systems: A Meta-Analysis of Methods, Applications, and Usability Challenges. Healthcare, 13(17), (2025) 2154. https://doi.org/10.3390/healthcare13172154