Abstract
Gradient-boosting classifiers now underpin most screening tools proposed for postpartum mental health, but published comparisons rarely separate what the modelling pipeline contributes from what the clinical construct itself contributes. Here we benchmark XGBoost, CatBoost, and LightGBM inside a single leakage-controlled pipeline, asking whether performance gaps trace back to identifiable structural properties rather than inconsistent data handling. A public postpartum survey dataset of 1,503 records was split 80:20 into development and test sets, with encoding and imputation fitted on the development set alone, and each classifier tuned via stratified five-fold cross-validation over a shared search space, isolating architecture as the sole variable. LightGBM's histogram-based, leaf-wise growth produced the strongest held-out discrimination (accuracy 88.4%, weighted F1 88.4%, balanced accuracy 87.8%, macro F1 88.0%), ahead of XGBoost (86.0%) and CatBoost (82.4%), with its advantage concentrated in the Moderate and High categories (81 of 92 High-category records correctly classified, versus 80 and 76). The selected classifier was integrated into a four-layer architecture, preprocessing, predictive inference, dialogue management, and deterministic recommendation, keeping the model's output separate from user-facing text and independently auditable. We report the full search space and the structural mechanisms plausibly explaining the ranking, and state plainly that training time, memory, and inference latency were not measured, naming this as required future work under TRIPOD+AI. The results demonstrate a reproducible method for comparing boosting architectures on a fixed screening task, they do not establish diagnostic validity, calibrated suicide-risk estimation, or clinical readiness, which require external validation beyond this internal benchmark.
Keywords
Gradient Boosting, Xgboost, Lightgbm, Catboost, Hyperparameter Optimization, Postpartum Mental Health, Algorithmic Architecture, Computational Performance, Safety-Aware Artificial Intelligence, Conversational Agent,Downloads
References
- Y. Yildiz, A. Kalayci, Gradient Boosting Decision Trees on Medical Diagnosis over Tabular Data. IEEE International Conference on AI and Data Analytics (ICAD), IEEE, USA. https://doi.org/10.1109/ICAD65464.2025.11114069
- R. Zhang, Y. Liu, Z. Zhang, R. Luo, B. Lv. Interpretable Machine Learning Model for Predicting Postpartum Depression: Retrospective Study. JMIR Medical Informatics, 13, (2025) e58649. https://doi.org/10.2196/58649
- W. Qi, Y. Wang, Y. Wang, S. Huang, C. Li, H. Jin, J. Zuo, X. Cui, Z. Wei, Q. Guo, J. Hu. Prediction of Postpartum Depression in Women: Development and Validation of Multiple Machine Learning Models. Journal of Translational Medicine, 23, (2025) 291. https://doi.org/10.1186/s12967-025-06289-6
- Z. Ma, M. Horvath, D.M. Stamilio, K. Sekyere, M.N. Gurcan. Building a Machine Learning Model to Predict Postpartum Depression from Electronic Health Records in a Tertiary Care Setting. Journal of Clinical Medicine, 14, (2025) 6644. https://doi.org/10.3390/jcm14186644
- X. Huang, L. Zhang, C. Zhang, J. Li, C. Li. Postpartum Depression Risk Prediction Using Explainable Machine Learning Algorithms. Frontiers of Medicine, 12, (2025) 1565374. http://dx.doi.org/10.3389/fmed.2025.1565374
- H. Kim. Predictive Analysis of Postpartum Depression Using Machine Learning. Healthcare, 13, (2025) 897. https://doi.org/10.3390/healthcare13080897
- C. Sona Ajay, S. Juliet, A. Alex. Machine Learning and Survival Analysis Models for Postpartum Depression: A Comprehensive Risk Factor Analysis. In Proceedings of the 2024 10th International Conference on Advanced Computing and Communication Systems (ICACCS), India. https://doi.org/10.1109/ICACCS60874.2024.10717207
- S. García-Méndez, F. de Arriba-Pérez. Detecting and Explaining Postpartum Depression in Real-Time with Generative Artificial Intelligence. Applied Artificial Intelligence, 39, (2025) 2515063. https://doi.org/10.1080/08839514.2025.2515063
- V.S.F. Gonçalves, V.R. de Carvalho. A Review of Interpretability Methods for Gradient Boosting Decision Trees. Journal of the Brazilian Computer Society, 31, (2025) 639–653. https://doi.org/10.5753/jbcs.2025.5324
- H. Semmelrock, T. Ross-Hellauer, S. Kopeinik, D. Theiler, A. Haberl, S. Thalmann, D. Kowald. Reproducibility in Machine-Learning-Based Research: Overview, Barriers, and Drivers. AI Magazine, 46(2), (2025) e70002. https://doi.org/10.1002/aaai.70002
- M. Alkhateeb, A. Nayeem, A. Ahmed, M. Alsahli, J. Sheikh, A. Abd-Alrazaq. AI for Detecting and Predicting Postpartum Depression: Scoping Review. Journal of Medical Internet Research, 28, (2026) e77376. https://doi.org/10.2196/77376
- U.S. Anaduaka, A.O. Oladosu, S. Katsande, C.S. Frempong, S. Awuku-Amador. Leveraging Artificial Intelligence in the Prediction, Diagnosis and Treatment of Depression and Anxiety among Perinatal Women in Low- and Middle-Income Countries: A Systematic Review. BMJ Mental Health, 28, (2025) e301445. https://doi.org/10.1136/bmjment-2024-301445
- J. Xia, C. Chen, X. Lu, T. Zhang, T. Wang, Q. Wang, Q. Zhou, Artificial Intelligence-Oriented Predictive Model for the Risk of Postpartum Depression: A Systematic Review. Front. Public Health, 13, (2025) 1631705. https://doi.org/10.3389/fpubh.2025.1631705
- L. Sibbald, M.I. van den Heuvel, M.R. Haas, C.J. van Lissa, H.J.A. van Bakel, J. Jongerling, L. P. Hulsbosch, L. Muskens, M. G. B. M. Boekhorst, I. Schwabe. Identifying Prenatal Risk Factors of Postpartum Depression with Machine Learning. Scientific Reports, 15, (2025) 34610. https://doi.org/10.1038/s41598-025-18204-6
- Y. Xie, H. Zheng, W. Gan, C. Su, M. Shams, J. Yang. The Performance of Machine Learning Models in Predicting Postpartum Depression: A Meta-Analysis and Systematic Review. Journal of Reproductive and Infant Psychology, 43(5), (2025) 1093–1110. https://doi.org/10.1080/02646838.2025.2517103
- L. Ou, Q. Shen, M. Xiao, W. Wang, T. He, B. Wang. Prevalence of Co-Morbid Anxiety and Depression in Pregnancy and Postpartum: A Systematic Review and Meta-Analysis. Psychological Medicine, 55, (2025) e93. https://doi.org/10.1017/S0033291725000601
- K. Zivin, C. Zhong, A. Rodríguez-Putnam, E. Spring, Q. Cai, A. Miller, L. J. Johns, V. A. Kalesnikava, A. Courant, B. Mezuk. Suicide Mortality during the Perinatal Period. JAMA Netw Open, 7, (2024) e2418887. https://doi.org/10.1001/jamanetworkopen.2024.18887
- H. Yu, Q. Shen, E. Bränn, Y. Yang, A.S. Oberg, U.A. Valdimarsdóttir, D. Lu. Perinatal Depression and Risk of Suicidal Behavior. JAMA Netw. Open, 7, (2024) e2350897. https://doi.org/10.1001/jamanetworkopen.2023.50897
- S. Goldman-Mellor, M. Olfson, A. Gemmill, C. Margerison. Incidence and Risk Factors for Suicide Attempt during Pregnancy and the Postpartum Period. The Journal of Clinical Psychiatry, 86, (2025) 24m15633. https://doi.org/10.4088/JCP.24m15633
- J. M. Martínez-Ramírez, R.A. Peinado-Molina, A. Hernández-Martínez, J.M. Martínez-Galiano. Predictive Models of Suicidal Ideation Risk in Perinatal-Stage Women Based on Sociodemographic and Clinical Data. Frontiers in Artificial Intelligence Medicine and Public Health, 9, (2026) 1774453. https://doi.org/10.3389/frai.2026.1774453
- Y. Hua, S. Siddals, Z. Ma, I. Galatzer-Levy, W. Xia, C. Hau, H. Na, M. Flathers, J. Linardon, C. Ayubcha, J. Torous. Charting the Evolution of Artificial Intelligence Mental Health Chatbots from Rule-Based Systems to Large Language Models: A Systematic Review. World Psychiatry, 24(3), (2025) 383–394. https://doi.org/10.1002/wps.21352
- A. Arnaiz-Rodriguez, M. Baidal, E. Derner, J. Layton Annable, M. Ball, M. Ince, E. Perez Vallejos, N. Oliver, Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models. JMIR Mental Health, 13(1), (2026) e88435. https://doi.org/10.2196/88435
- A.-R. Wrightson-Hester, G. Anderson, J. Dunstan, P. M. McEvoy, C.J. Sutton, B. Myers, S.J. Egan, S.J. Tai, M. Johnston-Hollitt, W. Chen, T. Gedeon, J.C. Moullin, W. Mansell, A Rule-Based Conversational Agent for Mental Health and Well-Being in Young People: Formative Case Series during the Rise of Generative AI. JMIR Formative Research, 9, (2025) e69841. https://doi.org/10.2196/69841
- G.S. Collins, K.G.M. Moons, P. Dhiman, R.D. Riley, A.L. Beam, B. van Calster, M. Ghassemi, X. Liu, J.B. Reitsma, M. van Smeden, A.L. Boulesteix, TRIPOD+AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods. BMJ, 385, (2024) e078378. https://doi.org/10.1136/bmj-2023-078378
- C. Meaney, X. Wang, J. Guan, T.A. Stukel, Comparison of Methods for Tuning Machine Learning Model Hyper-Parameters: with Application to Predicting High-Need High-Cost Health Care Users. BMC Medical Research Methodology, 25(1), (2025) 134. https://doi.org/10.1186/s12874-025-02561-x
- V. Gupta, S. Tripathi, D. Singh, A. Bansal, (2024) Predictive Algorithms for Early Postpartum Depression Detection: CatBoost vs. LightGBM. In Proceedings of the 2024 11th International Conference on Reliability, Infocom Technologies and Optimization (ICRITO), IEEE, Noida, India. https://doi.org/10.1109/ICRITO61523.2024.10522300
- N. Hagatulah, E. Bränn, A.S. Oberg, U.A. Valdimarsdóttir, D. Lu, Q. Shen, Perinatal Depression and Risk of Mortality: Nationwide, Register Based Study in Sweden. BMJ, 384, (2024) e075462. https://doi.org/10.1136/bmj-2023-075462
- Z. Kaminsky, R. J. McQuaid, K. G. Hellemans, Z. R. Patterson, M. Saad, R. L. Gabrys, T. Kendzerska, A. Abizaid, R. Robillard, Machine Learning-Based Suicide Risk Prediction Model for Suicidal Trajectory on Social Media Following Suicidal Mentions: Independent Algorithm Validation. Journal of Medical Internet Research, 26, (2024) e49927. https://doi.org/10.2196/49927
- L. Liu, Z. Li, Y. Hu, C. Li, S. He, S. Zhang, J. Gao, H. Zhu, G. Huang, Predictive Performance of Machine Learning for Suicide in Adolescents: Systematic Review and Meta-Analysis. Journal of Medical Internet Research, 27, (2025) e73052. https://doi.org/10.2196/73052
- R. Li, Y. Yue, X. Gu, L. Xiong, M. Luo, L. Li, Risk Prediction Models for Adolescent Suicide: A Systematic Review and Meta-Analysis. Psychiatry Research, 347, (2025) 116405. https://doi.org/10.1016/j.psychres.2025.116405
- L. Deng, K. Lu, H. Hu, An Interpretable LightGBM Model for Predicting Coronary Heart Disease: Enhancing Clinical Decision-Making with Machine Learning. PLoS one, 20(9), (2025) e0330377. https://doi.org/10.1371/journal.pone.0330377
- J. Xu, T. Chen, X. Fang, L. Xia, X. Pan, Prediction Model of Pressure Injury Occurrence in Diabetic Patients during ICU Hospitalization: XGBoost Machine Learning Model Can Be Interpreted Based on SHAP. Intensive and Critical Care Nursing, 83, (2024) 103715. https://doi.org/10.1016/j.iccn.2024.103715
- M.R. Khurshid, S. Manzoor, T. Sadiq, L. Hussain, M.S. Khan, A.K. Dutta, Unveiling Diabetes Onset: Optimized XGBoost with Bayesian Optimization for Enhanced Prediction. PLoS ONE, 20, (2025) e0310218. https://doi.org/10.1371/journal.pone.0310218
- C. Yang, E.A. Fridgeirsson, J.A. Kors, J.M. Reps, P.R. Rijnbeek, Impact of Random Oversampling and Random Undersampling on the Performance of Prediction Models Developed Using Observational Health Data. Journal of Big Data, 11(1), (2024) 7. https://doi.org/10.1186/s40537-023-00857-7
- J. Digitale, D. Franzon, M.J. Pletcher, C.E. McCulloch, E.D. Gennatas, Methods for Addressing Missingness in Electronic Health Record Data for Clinical Prediction Models: Comparative Evaluation. JMIR Medical Informatics, 13(1), (2025) e79307. https://doi.org/10.2196/79307
- R. Sanjeewa, R. Iyer, P. Apputhurai, N. Wickramasinghe, D. Meyer, Empathic Conversational Agent Platform Designs and Their Evaluation in the Context of Mental Health: Systematic Review. JMIR Mental Health, 11, (2024) e58974. https://doi.org/10.2196/58974
- W. Pichowicz, M. Kotas, P. Piotrowski, Performance of Mental Health Chatbot Agents in Detecting and Managing Suicidal Ideation. Scientific Reports, 15(1), (2025) 31652. https://doi.org/10.1038/s41598-025-17242-4
- A. Palmer, D. Schwan, Digital Mental Health Tools and AI Therapy Chatbots: A Balanced Approach to Regulation. Hastings Center Report, 55(3), (2025) 15–29. https://doi.org/10.1002/hast.4979
- Y. Jin, J. Liu, P. Li, B. Wang, Y. Yan, H. Zhang, C. Ni, J. Wang, Y. Li, Y. Bu, Y. Wang, The Applications of Large Language Models in Mental Health: Scoping Review. Journal of Medical Internet Research, 27, (2025) e69284. https://doi.org/10.2196/69284
- Y. Hua, H. Na, Z. Li, et al., A Scoping Review of Large Language Models for Generative Tasks in Mental Health Care. npj digital medicine, 8, (2025) 230. https://doi.org/10.1038/s41746-025-01611-4
- Stolltho, Postpartum Depression Classification [Kaggle Notebook and Associated Input Data]. Kaggle, n.d. Available online: https://www.kaggle.com/code/stolltho/postpartum-depression-classification/input
- E. Zain, Y. Watanabe, S. Takabayashi, L. Por, S. Fujita, S. Moriyama, A. Honma, N. Fukui, S. Boku, Psychometric Evaluation of the Japanese Edinburgh Postnatal Depression Scale for Screening Postpartum Anxiety. Frontiers in Psychiatry, 16, (2025) 1659497. https://doi.org/10.3389/fpsyt.2025.1659497
- L. Liou, E. Scott, P. Parchure, Y. Ouyang, N. Egorova, R. Freeman, I. S. Hofer, G. N. Nadkarni, P. Timsina, A. Kia, M. A. Levin, Assessing Calibration and Bias of a Deployed Machine Learning Malnutrition Prediction Model within a Large Healthcare System. npj Digit. Med., 7, (2024) 141. https://doi.org/10.1038/s41746-024-01141-5
- I. Partheniadis, P. Talimtzi, A. Nikolakopoulou, A.-B. Haidich, Machine Learning-Based COVID-19 Prognostic Models Lag Behind in Reporting Quality: Findings from a TRIPOD/TRIPOD+AI Systematic Review. Diagn. Progn. Res., 10, (2026) 3. https://doi.org/10.1186/s41512-026-00218-x
- A. J. Millner, M. D. Lee, M. K. Nock, Single-Item Measurement of Suicidal Behaviors: Validity and Consequences of Misclassification. PLoS ONE, 10, (2015) e0141606. https://doi.org/10.1371/journal.pone.0141606
- S.E. Davis, C. Dorn, D.J. Park, M.E. Matheny, Emerging Algorithmic Bias: Fairness Drift as the Next Dimension of Model Maintenance and Sustainability. Journal of the American Medical Informatics Association, 32, (2025) 845–854. https://doi.org/10.1093/jamia/ocaf039
- K.G.M. Moons, J.A.A. Damen, T. Kaul, L. Hooft, C. Andaur Navarro, P. Dhiman, A.L. Beam, B. Van Calster, L.A. Celi, S. Denaxas, A.K. Denniston, M. Ghassemi, G. Heinze, A.P. Kengne, L. Maier-Hein, X. Liu, P. Logullo, M.D McCradden, N. Liu, L. Oakden-Rayner, ,K. Singh, D.S Ting, ,L. Wynants, ,B. Yang, ,J.B Reitsma, ,R.D Riley, G.S Collins, M.V. Smeden, PROBAST+AI: An Updated Quality, Risk of Bias, and Applicability Assessment Tool for Prediction Models Using Regression or Artificial Intelligence Methods. BMJ, 388, (2025) e082505. https://doi.org/10.1136/bmj-2024-082505
- Q. Abbas, W. Jeong, S.W. Lee, Explainable AI in Clinical Decision Support Systems: A Meta-Analysis of Methods, Applications, and Usability Challenges. Healthcare, 13(17), (2025) 2154. https://doi.org/10.3390/healthcare13172154
Articles

