Khalil MAM, Sadagah NM, Tan J, Al-Qurashi SH. Predicting delayed graft function after kidney transplant: Do complex models help compared to standard statistics? World J Nephrol 2026; 15(3): 120300 [DOI: 10.5527/wjn.120300]
Corresponding Author of This Article
Muhammad Abdul Mabood Khalil, FRCP, Renal Diseases and Transplantation Center, King Fahad Armed Forces Hospital, Al Kurnaysh Br Road, Al Andalus, Jeddah 23311, Makkah al Mukarramah, Saudi Arabia. doctorkhalil1975@hotmail.com
Research Domain of This Article
Transplantation
Article-Type of This Article
editorial
Open-Access Policy of This Article
This article is an open-access article which was selected by an in-house editor and fully peer-reviewed by external reviewers. It is distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/
Baishideng Publishing Group Inc, 7041 Koll Center Parkway, Suite 160, Pleasanton, CA 94566, USA
Share the Article
Khalil MAM, Sadagah NM, Tan J, Al-Qurashi SH. Predicting delayed graft function after kidney transplant: Do complex models help compared to standard statistics? World J Nephrol 2026; 15(3): 120300 [DOI: 10.5527/wjn.120300]
Muhammad Abdul Mabood Khalil, Nihal Mohammed Sadagah, Salem H Al-Qurashi, Renal Diseases and Transplantation Center, King Fahad Armed Forces Hospital, Jeddah 23311, Makkah al Mukarramah, Saudi Arabia
Jackson Tan, Department of Nephrology, RIPAS Hospital Brunei Darussalam, Bander Seri Begawan BA1712, Brunei Darussalam
Author contributions: Khalil MAM, Sadagah NM, Al-Qurashi SH, and Tan J planned and designed the outline of the manuscript; Khalil MAM wrote the manuscript; Sadagah NM, Tan J, and Al-Qurashi SH helped in the literature search and supported in writing; all authors read and agreed to the final manuscript.
Conflict-of-interest statement: All authors declare that they have no conflict of interest to disclose.
Corresponding author: Muhammad Abdul Mabood Khalil, FRCP, Renal Diseases and Transplantation Center, King Fahad Armed Forces Hospital, Al Kurnaysh Br Road, Al Andalus, Jeddah 23311, Makkah al Mukarramah, Saudi Arabia. doctorkhalil1975@hotmail.com
Received: February 26, 2026 Revised: March 31, 2026 Accepted: April 16, 2026 Published online: September 25, 2026 Processing time: 168 Days and 5.8 Hours
Abstract
Both traditional statistics, such as the logistic regression (LR) model, and machine learning (ML) have strengths and limitations for predicting outcomes after kidney transplantation. The LR model is simple, interpretable, and reliable with small datasets. ML can capture complex, nonlinear patterns and manage many variables, but it needs larger, high-quality datasets to reach its full potential. In the recent issue of World Journal of Nephrology, Salgado et al compared six ML models with the LR model using donor, transplant, and recipient data from 523 deceased-donor kidney transplants. Surprisingly, ML models only slightly outperformed the LR model, and overall predictive performance remained modest, especially for identifying patients who developed delayed graft function. These results emphasize that dataset size, completeness, and relevant clinical variables may be more important than algorithm complexity. Future work should focus on improving data quality and developing models that are both accurate and clinically interpretable.
Core Tip: Delayed graft function significantly impacts kidney transplant outcomes, yet predicting it remains challenging. Recent evidence shows that machine learning (ML) models offer only modest improvements over the traditional logistic regression model when the dataset is of limited quality. High-quality, comprehensive data and interpretable models are critical for accurate risk stratification. Integrating ML with transparent statistical approaches may optimize predictive performance and support clinically meaningful decision-making in transplantation.
Citation: Khalil MAM, Sadagah NM, Tan J, Al-Qurashi SH. Predicting delayed graft function after kidney transplant: Do complex models help compared to standard statistics? World J Nephrol 2026; 15(3): 120300
This editorial refers to “Prediction of graft outcomes after kidney transplantation: When standard statistics compare to machine learning techniques” by Salgado et al, 2026; https://doi.org/10.5527/wjn.v15.i1.116879.
INTRODUCTION
In the recent issue of World Journal of Nephrology, Salgado et al[1] present a methodologically rigorous study on delayed graft function (DGF) in kidney transplantation, evaluating traditional statistical methods, including the logistic regression (LR) model, as well as assessing machine learning (ML) techniques for predictive performance. Their study highlights how donor, recipient, and procedural variables interact to influence post-transplant outcomes and underscores the potential and limitations of artificial intelligence in this setting.
DGF affects 20%-50% of deceased donor kidney transplant recipients[2]. It is associated with an increased risk of acute rejection and reduced graft survival[3], prolonged hospitalization and higher healthcare costs[4], slower recovery of renal function[5], and may contribute to long-term allograft dysfunction[2]. Given its substantial impact on both short- and long-term outcomes, early identification of patients at risk for DGF is crucial.
Accurate risk stratification helps transplant teams individualize perioperative management and also allows efforts to reduce cold ischemia time and to implement preventive measures to protect the graft[6]. Prompt recognition and proactive management may reduce complications such as prolonged dialysis, extended hospital stays, and early rejection episodes, ultimately improving graft and patient survival. Reliable predictive tools are, therefore, not merely academic instruments but practical necessities in modern kidney transplantation.
The development of DGF is multifactorial, with donor-related characteristics including advanced age, preexisting renal dysfunction, and extended criteria donation, significantly increasing risk[7]. Recipient-related factors such as obesity[8], long-term dialysis, increased human leukocyte antigen mismatches, and diabetes[9] may impair post-transplant recovery. Procedural variables, including prolonged cold ischemia time and surgical complications, can further exacerbate ischemia-reperfusion injury[10]. Recognizing these interacting risk domains before transplantation allows clinicians to optimize donor selection, refine perioperative care, and apply predictive models in a clinically meaningful way.
In this context, the development of reliable predictive models for DGF has become a major focus of transplant research. Salgado et al[1] provide a timely, methodologically rigorous comparison of traditional statistical approaches, such as the LR model, with ML techniques for DGF prediction. By systematically evaluating donor, recipient, and transplant-related variables, their study offers valuable insight into the strengths and limitations of artificial intelligence in this setting. Importantly, their findings contribute to an ongoing discussion: Does increasing algorithmic complexity translate into meaningful clinical improvement?
Collectively, these considerations underscore the continued need for accurate, clinically applicable tools to predict DGF. While traditional statistical methods remain foundational, emerging ML techniques offer the potential to integrate multiple interacting variables into more comprehensive risk assessments[11]. Understanding where each approach adds value is essential to improving outcomes for kidney transplant recipients.
Unlike many previous reviews that mainly compare the predictive performance of ML and traditional statistical models, this editorial takes a slightly different perspective. We emphasize that the usefulness of predictive models for DGF depends not only on the complexity of the algorithm but also on the quality of the dataset, the completeness of clinically relevant variables, and the model’s interpretability. In this context, we discuss the importance of integrating ML with traditional statistical approaches to develop prediction tools that are both accurate and clinically applicable, particularly in real-world transplant settings where data limitations are common.
MACHINE LEARNING IN TRANSPLANTATION: DREAM OR REALITY?
ML learning has generated considerable enthusiasm in transplantation research, promising improved predictive accuracy and more personalized patient care. At its core, ML leverages algorithms that identify patterns in large datasets[11]. By analyzing extensive donor, recipient, and transplant-related variables, ML models can detect complex and potentially nonlinear associations with post-transplant outcomes[12].
Beyond DGF prediction, ML has been applied to forecasting graft survival[13], optimizing donor-recipient matching and organ allocation[14], stratifying immunological risk[15], detecting early graft dysfunction[16], tailoring immunosuppressive therapy[17], and estimating post-transplant complications[18]. These expanding applications illustrate the transformative potential of data-driven approaches in transplant medicine.
However, the promise of ML must be interpreted with appropriate caution as model performance depends heavily on data quality, completeness, feature engineering, and rigorous validation. A recent large cohort study of 1857 deceased-donor kidney transplant recipients applied a random forest model to predict DGF. The model demonstrated reasonable discrimination, with donation after circulatory death, recipient body mass index, and donor age showing the highest feature importance. Nevertheless, the authors emphasized that a direct comparison with the conventional LR model is ongoing, noting that ML is promising but not yet proven superior[19].
Thus, while ML undoubtedly expands methodological possibilities, its true clinical value must be critically evaluated alongside traditional statistical approaches rather than assumed a priori.
TRADITIONAL STATISTICS VS MACHINE LEARNING: COMPETITION OR COMPLEMENT?
As summarized in Table 1, across multiple studies, ML models generally achieved slightly higher performance metrics than LR models, though differences were modest and often not statistically significant.
Table 1 Studies comparing machine learning and logistic regression for predicting delayed graft function.
ML models showed modestly higher AUROC/accuracy for some predictor sets, but overall differences between ML and LR were not statistically significant, and both approaches showed limited sensitivity
Salgado et al[1] provide important clarity in this debate. Their analysis demonstrated that ML models did not significantly outperform traditional LR models in predicting DGF, and key performance indicators, including area under the curve (AUC), accuracy, sensitivity, and specificity, were largely comparable across approaches.
Importantly, although overall discrimination appeared similar, sensitivity for predicting DGF remained modest across most models. From a clinical standpoint, this imbalance between sensitivity and specificity is highly relevant, as failure to reliably identify high-risk patients may limit the practical utility of predictive tools. Statistical equivalence does not necessarily translate into clinical adequacy, particularly when early intervention depends on reliably identifying vulnerable recipients.
These findings align with prior research. In a cohort of 497 deceased-donor kidney transplants, Decruyenaere et al[20] compared nine predictive models, including LR model, linear discriminant analysis (LDA), quadratic discriminant analysis (QDA), multiple support vector machine (SVM) variants, decision tree, random forest, and stochastic gradient boosting. Overall, none of the pairwise area under the receiver operating characteristic curve (AUROC) comparisons were statistically significant except for linear SVM vs LR model, suggesting that ML models generally offer modest incremental improvement over traditional statistics in structured datasets[20]; the observed incidence of DGF was 12.5%. Only linear SVM slightly outperformed LR model (AUROC: 84.3% vs 81.7%), while other models, including LDA, radial SVM, and QDA, showed discriminative capacities similar to LR model. Comparable observations have been reported in other areas of clinical prediction, where advanced ML approaches often provide only modest improvements over simpler methods[21]. Together, these findings emphasize a critical lesson that algorithmic complexity alone does not guarantee superior predictive accuracy.
Moreover, dataset quality and completeness remain central determinants of performance. In many transplant registries, including the cohort analyzed by Salgado et al[1], missing perioperative variables and limited representation of donor hemodynamics or biomarker data reflect broader systemic challenges rather than isolated study limitations[22]. When key predictors are absent or incompletely captured, even sophisticated ML algorithms may be unable to fully exploit nonlinear relationships, thereby narrowing any potential performance gap with LR model[21].
In such contexts, the transparency and interpretability of the LR model provide substantial advantages[23]. Rather than framing ML and LR models as competitors, current evidence suggests they are better viewed as complementary tools whose utility depends on dataset characteristics, clinical objectives, and interpretability requirements.
INTERPRETABILITY AND CLINICAL TRUST
A persistent limitation of many ML models is their “black box” nature, which can hinder clinician trust and adoption[24]. While these algorithms may detect complex interactions, the lack of transparency makes it difficult to understand how predictions are generated.
In contrast, traditional statistical models such as the LR model provide explicit and interpretable relationships between predictors and outcomes[23]. This transparency allows clinicians to verify whether predictions align with established clinical knowledge, detect potential inconsistencies, and understand the boundaries of model applicability.
Although explainability techniques such as Shapley values[25], partial dependence plots (PDPs)[26], and feature importance metrics can improve insight into ML models, they do not fully eliminate uncertainty or guarantee clinician confidence. Moreover, reliance on non-transparent models can create challenges when predictive performance is imperfect or when unexpected outcomes occur[27].
Interpretability is therefore not merely a methodological preference—it is central to patient safety, shared decision-making, and responsible integration of predictive analytics into clinical practice[28].
WAY FORWARD: INTEGRATING MACHINE LEARNING AND TRADITIONAL METHODS IN DGF PREDICTION
Salgado et al[1] challenge the assumption that increasing algorithmic complexity automatically yields superior predictive accuracy with their findings underscoring that data quality, cohort size, and feature completeness are as influential as model architecture.
Recent studies further illustrate the expanding role of ML in DGF prediction. Konieczny et al[29] applied random forest classifiers and multilayer perceptron to deceased donor transplants, identifying donor and recipient body mass index, donor estimated glomerular filtration rate, and weight mismatch as key predictors with high accuracy and AUC values. In turn, Liu et al[30] extended this approach to pediatric recipients, developing an integrated ML-based DGF risk score validated through decision curve analysis.
He et al[31] demonstrated the utility of deep learning in a cohort of 670 adult recipients, achieving balanced sensitivity and specificity, while Tirasattayapitak et al[32] integrated clinical and histopathological features from time-zero biopsies and found that Extreme Gradient Boosting outperformed neural networks and random forests, although the LR model remained highly interpretable. In a large registry-based study of over 55000 recipients, Kawakita et al[33] demonstrated that ML modestly outperformed the LR model, highlighting scalability in real-world datasets.
Collectively, these studies emphasize that predictive performance depends fundamentally on high-quality, comprehensive data. Equally important are calibration and clinical utility assessments, such as decision curve analysis, to determine whether statistical discrimination translates into meaningful bedside benefit[34,35].
ML requires accurate input from large, heterogeneous populations spanning continents, countries, and centers. A multilateralism involving registry participation and internal research collaboration is essential to enhance understanding and accuracy of the algorithms.
Future research should, therefore, move beyond algorithm comparison alone[21]. The critical question is not which model achieves a marginally higher AUC, but whether predictive tools meaningfully alter clinical decision-making, improve resource allocation, and enhance graft and patient outcomes.
Hybrid or interpretable modeling strategies that maintain transparency while capturing complex patterns may represent the most promising path forward. As such, multicenter registries, standardized definitions of DGF, prospective validation, and harmonized data collection will be essential to realizing this vision.
To improve clinical acceptability, future research should make ML models more explainable. SHapley Additive exPlanations values can show each feature’s contribution to predictions[36], and PDPs can show how predictors affect outcomes across patients[37]. These methods can make DGF prediction models more transparent and help clinicians trust the results. Prospective studies should also test how interpretable ML outputs influence real-world decisions, such as early monitoring, adjusting immunosuppression, or donor selection. Embedding explainability can make ML models useful in everyday kidney transplant care.
CONCLUSION
The predictive landscape of DGF is shaped as much by data integrity and completeness as by algorithmic sophistication. ML offers powerful tools to uncover complex interactions and refine risk stratification, while the traditional LR regression model ensures interpretability, transparency, and clinician confidence. The future of DGF prediction lies in combining the complementary strengths of both approaches. Progress will, however, require richer datasets, standardized reporting, rigorous external validation, and careful assessment of clinical impact. By integrating these strategies, predictive modeling can move from theory to practice, ultimately improving graft outcomes and patient care in kidney transplantation.
Fan B, Schürch M, Tian Y, Mallone A, Frischknecht L, Koller M, Van Delden C, Leichtle A, Golshayan D, Villard J, Schachtner T, Sidler D, Schaub S, Nilsson J, Krauthammer M; Swiss Transplant Cohort Study. Enhancing post-kidney transplant prognostication: an interpretable machine learning approach for longitudinal outcome prediction.NPJ Digit Med. 2025;8:684.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 3][Cited by in RCA: 3][Article Influence: 3.0][Reference Citation Analysis (0)]
Kim J, Minkovich M, Azafar G, Brilian M, Li Y, Debuono S, Panesar J, Famure O. Predicting Delayed Graft Function After Kidney Transplantation: A New Look at an Old Problem.Am J Transplant. 2025;25:S871.
[PubMed] [DOI] [Full Text]
Liu XY, Feng RT, Feng WX, Jiang WW, Chen JA, Zhong GL, Chen CW, Li ZJ, Zeng JD, Liu D, Zhou S, Hu JM, Liao GR, Liao J, Guo ZF, Li YZ, Yang SQ, Li SC, Chen H, Guo Y, Li M, Fan LP, Yan HY, Chen JR, Li LY, Liu YG. An integrated machine learning model enhances delayed graft function prediction in pediatric renal transplantation from deceased donors.BMC Med. 2024;22:407.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 9][Cited by in RCA: 7][Article Influence: 3.5][Reference Citation Analysis (1)]
Tirasattayapitak S, Ratanatharathorn C, Thotsiri S, Sutharattanapong N, Wiwattanathum P, Arpornsujaritkun N, Sirisopana K, Worawichawong S, Rostaing L, Kantachuvesiri S. Integrating Clinical and Histopathological Data to Predict Delayed Graft Function in Kidney Transplant Recipients Using Machine Learning Techniques.J Clin Med. 2024;13:7502.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 3][Cited by in RCA: 4][Article Influence: 2.0][Reference Citation Analysis (0)]