Su Y, Wang DX, Zhao YQ, Xing X. Bid farewell to single indicators: Machine learning models integrating multidimensional data lead thrombosis risk prediction into a new stage. World J Gastroenterol 2026; 32(34): 118337 [DOI: 10.3748/wjg.118337]
Corresponding Author of This Article
Xue Xing, Department of Clinical Laboratory, The Second Affiliated Hospital of Dalian Medical University, No. 467 Zhongshan Road, Dalian 116021, Liaoning Province, China. dyeyxx39198645@126.com
Research Domain of This Article
Gastroenterology & Hepatology
Article-Type of This Article
editorial
Open-Access Policy of This Article
This article is an open-access article which was selected by an in-house editor and fully peer-reviewed by external reviewers. It is distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/
Baishideng Publishing Group Inc, 7041 Koll Center Parkway, Suite 160, Pleasanton, CA 94566, USA
Share the Article
Su Y, Wang DX, Zhao YQ, Xing X. Bid farewell to single indicators: Machine learning models integrating multidimensional data lead thrombosis risk prediction into a new stage. World J Gastroenterol 2026; 32(34): 118337 [DOI: 10.3748/wjg.118337]
Ying Su, Dong-Xia Wang, Yi-Qun Zhao, Xue Xing, Department of Clinical Laboratory, The Second Affiliated Hospital of Dalian Medical University, Dalian 116021, Liaoning Province, China
Co-corresponding authors: Yi-Qun Zhao and Xue Xing.
Author contributions: Su Y and Wang DX contributed equally as co-first authors; Xing X designed the overall concept and outline of the manuscript; Su Y contributed to the discussion and design of the manuscript; Wang DX and Zhao YQ contributed to the writing and editing the manuscript, illustrations, and review of literature; Zhao YQ and Xing X contributed equally as co-corresponding authors. All authors approved the final version to publish.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
Corresponding author: Xue Xing, Department of Clinical Laboratory, The Second Affiliated Hospital of Dalian Medical University, No. 467 Zhongshan Road, Dalian 116021, Liaoning Province, China. dyeyxx39198645@126.com
Received: December 30, 2025 Revised: February 2, 2026 Accepted: February 12, 2026 Published online: September 14, 2026 Processing time: 232 Days and 14.8 Hours
Abstract
This editorial reviews a multicenter study published by Lu et al in the World Journal of Gastroenterology. The study innovatively applied five machine learning models to predict the risk of thromboembolic events in patients with non-variceal gastrointestinal bleeding. The research found that the simplified model constructed based on 10 key clinical variables (including D-dimer, age, anticoagulant history, etc.) significantly outperformed the traditional single indicator of D-dimer in predictive performance. Among them, the classification boosting algorithm model showed the best discrimination and calibration in external validation. This commentary delves into the milestone significance of the study in the field of gastroenterology risk stratification, pointing out that it marks a paradigm shift from reliance on a single biomarker to the integration of multidimensional clinical data. We further analyze the potential advantages of machine learning models compared to traditional scoring systems (such as the Padua score) and the challenges faced in clinical integration, emphasizing that it provides a powerful decision support tool for achieving personalized and precise thrombus prevention management for nonvariceal gastrointestinal bleeding patients, and looks forward to future research exploring prospective validation, algorithm interpretability, and clinical workflow integration.
Core Tip: This multicenter study demonstrates that machine learning models based on routine clinical data can effectively predict thrombotic risk in patients with nonvariceal gastrointestinal bleeding, outperforming the traditional D-dimer biomarker. The classification boosting algorithm model exhibited the best performance, aiding in the clinical identification of high-risk patients for early intervention while avoiding excessive monitoring in low-risk individuals, thereby advancing thromboprophylaxis strategies toward precision medicine.
Citation: Su Y, Wang DX, Zhao YQ, Xing X. Bid farewell to single indicators: Machine learning models integrating multidimensional data lead thrombosis risk prediction into a new stage. World J Gastroenterol 2026; 32(34): 118337
This editorial refers to “Application of machine learning models in predicting the risk of thromboembolic events in patients with nonvariceal gastrointestinal bleeding” by Lu et al, 2026; https://doi.org/10.3748/wjg.v32.i3.115527
INTRODUCTION
Nonvariceal gastrointestinal bleeding (NVGIB) is a critical emergency characterized by diverse clinical manifestations and complex etiologies, including common causes such as peptic ulcers, tumors, and rare vascular anomalies. It is associated with high mortality and a propensity for hemorrhagic shock and severe anemia, posing significant challenges in clinical management[1,2]. While advances in endoscopic hemostasis have improved bleeding management, the challenge of balancing hemorrhage control with thrombosis prevention persists[3].
Thromboembolism represents a major complication in NVGIB patients, substantially increasing both therapeutic complexity and mortality[4]. While the use of antiplatelet and anticoagulant agents can reduce thrombotic risk, it concurrently elevates the bleeding risk in this population[5,6]. Consequently, clinical decision-making often faces the dilemma of balancing hemorrhage control against thrombosis prevention.
Machine learning, as a data-driven intelligent analytical approach, is capable of handling complex nonlinear relationships among variables, thereby enabling precise individualized risk prediction. It holds considerable clinical application potential for predicting thrombotic disorders such as deep vein thrombosis, pulmonary embolism, and cerebral infarction (ischemic stroke)[7-10]. The recent study completed by Lu et al[11] represents the first systematic development and validation of machine learning prediction models specifically for the NVGIB population. This work is a significant contribution at the intersection of gastroenterology and artificial intelligence, with its model’s predictive efficacy markedly superior to that of traditional biomarkers, providing a powerful new tool for clinical decision-making.
PRECISE RISK IDENTIFICATION: CAPTURING KEY VARIABLES IN A SPECIFIC CONTEXT
The management of thromboembolic risk in patients with NVGIB consistently presents a clinical paradox: The urgent need to achieve hemostasis while simultaneously preventing thrombosis. The study by Lu et al[11] addresses this by employing SHapley Additive exPlanations analysis to identify ten predictive factors. This list includes not only established risk factors like age and prior thrombosis but also reveals variables previously overlooked in the NVGIB context. A notable example is “the use of hemostatic agents”, identified here as a significant risk factor, a finding that appears to contrast with some prior meta-analyses[12].
This discrepancy likely underscores the unique, dynamic imbalance between hemostasis and coagulation in NVGIB, a state often aggravated by active bleeding and hemodynamic instability. The strength of the machine learning model lies in its capacity to quantify these complex, nonlinear relationships from high-dimensional data, offering an analytical depth that surpasses traditional statistics and clinical intuition[13,14]. Consequently, the model demonstrated superior sensitivity and specificity, proving particularly advantageous for the early identification of high-risk patients.
While this data-driven approach optimizes resource allocation and provides a foundation for personalized treatment, the study’s retrospective design introduces inherent limitations. Data collected under non-controlled conditions carry risks of information and selection bias. For instance, more complete data from sicker or more frequently monitored patients may lead the model to learn patterns of care intensity rather than true biological risk. Strong performance on historical data does not guarantee efficacy in future patients, highlighting the critical importance of prospective validation in subsequent research.
TRANSCENDING TRADITIONAL SCORES: A COMPREHENSIVE PROFILE INTEGRATING MULTIDIMENSIONAL DATA
Clinical data for predicting thrombotic risk in NVGIB are typically heterogeneous, plagued by irrelevant features, missing values, and class imbalance[15,16]. This reality makes robust feature selection and data preprocessing critical for building reliable machine-learning models. The study by Lu et al[11] powerfully underscores the advantage of such models over traditional scoring systems like the Padua score[17]. Table 1 provides a comparative overview of key performance metrics and integration capabilities between the CatBoost model, the traditional D-dimer biomarker, and the Padua prediction score, highlighting the significant leap forward in risk prediction accuracy and individualization. While Padua score relies on limited, static parameters for venous thromboembolism screening, their model integrates dynamic, multidimensional data, including laboratory indices and treatment decisions (e.g., intensive care unit admission, hemostatic agent use), to generate a holistic, individualized risk profile capable of predicting a wider spectrum of thrombotic events[18]. This data-integration capability is vital for managing complex, acute conditions like NVGIB.
Table 1 Clinical applicability comparison (machine learning models vs traditional markers).
However, the study’s discussion of data completeness and processing remains underdeveloped. Although patients with “incomplete clinical data” were excluded, the handling of specific missing variables (e.g., sporadically performed lab tests) was not detailed. Inappropriate methods like simple deletion or mean imputation can compromise model reliability, and inadequate reporting on this front is a recognized flaw in predictive research[16].
Furthermore, variable selection was inherently limited to the available 42 features, potentially omitting important predictors such as genetic predisposition or dynamic inflammatory markers. Future work should expand clinically relevant variables and leverage advanced multimodal fusion techniques (e.g., multi-view learning) to achieve a more comprehensive pathophysiological representation, thereby significantly enhancing risk identification accuracy[19].
MODEL PERFORMANCE AND INNOVATION: BENCHMARKING A NEW STANDARD FOR CLINICAL RISK MANAGEMENT
Machine learning is increasingly applied to predict thromboembolic risk in patients with NVGIB[20-24]. In this context, where data are inherently heterogeneous and relationships complex, the choice and optimization of algorithms are critical to predictive performance. The innovation of the study by Lu et al[11] lies in its systematic approach to this challenge. The authors concurrently developed and compared five mainstream machine learning algorithms, including categorical boosting (CatBoost) and random forest, specifically for the NVGIB population. All models, in external validation, significantly outperformed the common clinical marker D-dimer. Notably, the CatBoost model distinguished itself by natively handling categorical features and resisting overfitting, achieving superior discrimination and calibration. This demonstrates its potential as a clinical decision-support tool. By validating this advanced framework within the urgent context of NVGIB, the study establishes a new performance benchmark and provides a valuable methodological reference for future research. To provide a holistic view of how machine learning models translate clinical data into actionable risk stratification, we constructed a conceptual application path diagram (Figure 1). This flowchart illustrates the sequential process from the ten identified clinical variables, through algorithmic processing, to the final output of risk stratification and corresponding clinical interventions.
Figure 1
Application path diagram of machine learning model in nonvariceal gastrointestinal bleeding thrombosis risk prediction.
Building on this evidence, the study enables a practical, phased risk-management strategy for current practice. Gastroenterologists should move beyond reliance on single indicators like D-dimer. Instead, clinical teams should systematically collect the ten key variables identified (e.g., dynamic D-dimer, albumin, detailed medication history) to form a comprehensive patient risk profile. Critical risk factors revealed by the model, such as “hemostatic drug use”, must be formally integrated into multidisciplinary discussions, especially when deciding on restarting anticoagulation. This factor should be acknowledged as a potential thrombotic risk modifier. Ultimately, applying this model facilitates personalized intervention strategies. By enabling dynamic risk stratification, it empowers clinicians to make more precise judgments in balancing bleeding and thrombosis risks, adjust treatment intensity proactively, and thereby improve overall management quality and patient prognosis.
CLINICAL CHALLENGES AND THE TRANSLATIONAL GAP
Although the study by Lu et al[11] has laid an important foundation, several practical chasms must be bridged for successful clinical translation. First, the level of evidence requires elevation. Conclusions from the current retrospective design must be validated through prospective, multicenter interventional trials to determine whether clinical decisions guided by the model’s risk stratification (e.g., initiating prophylactic anticoagulation in high-risk patients) can genuinely improve net clinical outcomes, reducing thrombotic events without significantly increasing bleeding risk[25]. As emphasized by Ramspek et al[26] in their comprehensive methodological framework, external validation is not merely a procedural formality but a critical step to assess a model’s reproducibility and transportability across different populations and settings[26-28]. The model’s generalizability also needs verification across broader populations and healthcare systems[29,30].
Second, a practical gap exists between prediction and action. The model outputs a probability, while clinical decisions are binary. Defining an “intervention threshold” that effectively identifies high-risk patients without leading to over-medicalization is itself a research endeavor requiring integration of local data. Furthermore, the model’s decision-making process remains opaque to clinicians, and building trust in such a “black-box” tool poses a significant barrier. It must be unequivocally established that machine learning-based clinical decision support systems are merely assistive tools; the ultimate decision-making authority and responsibility reside with the clinician, necessitating hospital protocols for use and audit. Finally, integration into clinical workflows is challenging. Deeply embedding the model into the hospital electronic health record system to enable automated risk assessment and alerts is far more than a technical interface; it is a systemic undertaking involving data standardization, system interoperability, and even the redesign of clinical pathways[31].
FUTURE DIRECTIONS: TOWARD DYNAMIC AND PRECISION PREVENTION
The core future direction lies in moving beyond single data dimensions to construct dynamic, individualized risk prediction models. This requires the integration of multimodal data from electronic health records, genomics, metabolomics, and real-time monitoring[32]. For instance, combining genetic susceptibility with real-time metabolic biomarkers could provide earlier insights into thrombotic tendencies, while continuous clinical and laboratory data offer dynamic inputs for the model to perceive disease progression. As Efthimiou et al[33] outline in their 13-step guide to prediction model development, this requires careful attention to data quality, handling of competing events, and iterative model updating to prevent calibration drift. By building such a multi-layered, real-time feedback intelligent decision-support ecosystem, we have the potential to fundamentally shift the management of thromboembolic risk in NVGIB patients from a “reactive” stance to a paradigm of “proactive and precision prevention”.
CONCLUSION
Lu et al[11] successfully apply machine learning to predict thrombotic risk in NVGIB patients, demonstrating performance superior to traditional methods. This work signals a shift in clinical paradigm, from focusing on isolated risk factors toward a holistic, quantitative risk assessment. With its robust performance, the CatBoost model emerges as a promising decision-support tool. Future efforts should focus on prospective, multinational validation and the seamless integration of such algorithms into clinical workflows. Ultimately, these tools aim to achieve the precise balance between bleeding and thrombosis, paving the way for individualized care in NVGIB.
Carballo F, Albillos A, Llamas P, Orive A, Redondo-Cerezo E, Rodríguez de Santiago E, Crespo J. Consensus document of the Spanish Society of Digestives Diseases and the Spanish Society of Thrombosis and Haemostasis on massive nonvariceal gastrointestinal bleeding and direct-acting oral anticoagulants.Rev Esp Enferm Dig. 2022;114:375-389.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 2][Cited by in RCA: 1][Article Influence: 0.3][Reference Citation Analysis (0)]
Yamaguchi D, Tominaga N, Miyahara K, Tsuruoka N, Sakata Y, Takeuchi Y, Matsunaga T, Hidaka H, Akutagawa T, Noda T, Ogata S, Tsunada S, Esaki M. Safety and Efficacy of the Noncessation Method of Antithrombotic Agents after Emergency Endoscopic Hemostasis in Patients with Nonvariceal Upper Gastrointestinal Bleeding: A Multicenter Pilot Study.Can J Gastroenterol Hepatol. 2021;2021:6672440.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 1][Cited by in RCA: 4][Article Influence: 0.8][Reference Citation Analysis (0)]
Boros E, Pintér J, Molontay R, Prószéky KG, Vörhendi N, Simon OA, Teutsch B, Pálinkás D, Frim L, Tari E, Gagyi EB, Szabó I, Hágendorn R, Vincze Á, Izbéki F, Abonyi-Tóth Z, Szentesi A, Vass V, Hegyi P, Erőss B. New machine-learning models outperform conventional risk assessment tools in Gastrointestinal bleeding.Sci Rep. 2025;15:6371.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 16][Cited by in RCA: 12][Article Influence: 12.0][Reference Citation Analysis (6)]
Creativity or innovation: Grade B, Grade B, Grade B
Scientific significance: Grade A, Grade A, Grade C
P-Reviewer: Qureshi H, PhD, Post Doctoral Researcher, Postdoctoral Fellow, Pakistan; Yan J, Head, Professor, China S-Editor: Wu S L-Editor: A P-Editor: Wang WB