Copyright: ©Author(s) 2026.
World J Stem Cells. Aug 26, 2026; 18(8): 121077
Published online Aug 26, 2026. doi: 10.4252/wjsc.121077
Published online Aug 26, 2026. doi: 10.4252/wjsc.121077
Table 1 Stem-cell relevance map linking each application area to stem-cell biological object, data modality, AI methodology, output, clinical/biological endpoint, and validation status
| Application | Stem-cell object | Data modality | AI method/model | Dataset/scale | Key performance metrics | Clinical/biological endpoint | Validation status | Ref. |
| A: Normal HSC biology | ||||||||
| HSC differentiation modeling | Normal HSCs, MPPs | scRNA-seq, scATAC-seq | VIA (Voyager) - lazy-teleporting MCMC trajectory inference | Human CD34+ hematopoiesis (multi-site) | F1 > 0.9 for rare lineage populations; robust across scRNA-seq and scATAC-seq | HSC differentiation dynamics; lineage bifurcation mapping | Cross-modal validation (scRNA-seq + scATAC-seq) | [5] |
| HSC aging - chromatin architecture | Young vs aged murine HSCs | 3D confocal DAPI chromatin imaging | ChromAgeNet - CNN on 3D nuclear images | 1229 HSC nuclei (551 young, 678 aged) | AUROC: 0.77 ± 0.03; accuracy 0.68 ± 0.05; outperforms handcrafted features (AUROC: 0.73) | Biological age prediction; epigenetic rejuvenation detection | 5-fold cross-validation; drug-treatment validation | [6] |
| HSC division dynamics | HSCs/progenitors - 3 age groups (young, mid-life, aged) | Single-cell gene expression | ANN | 9 stem/progenitor populations across 3 age groups | 89% accuracy (cell type + donor age); 96% accuracy (regenerative status) | Age-dependent self-renewal; niche sensitivity across lifespan | Cross-validation | [10] |
| HSC morphological classification | HSCs vs MPPs | Bright-field microscopy images | Deep CNN (morphology-based) | Steady-state bright-field image dataset | High accuracy distinguishing HSCs from MPPs; morphological features encode functional state; > 98% accuracy in leukemic BM context | Non-destructive functional state identification without molecular profiling | Steady-state + leukemic BM validation | [11] |
| HSC quiescence regulation | LTHSCs, STHSCs | scRNA-seq integrated with niche signals (TPO, SCF, ANGPT1) | Boolean network modeling (constraint programming + model checking) | Pseudotrajectory-derived gene expression states | Predicted stable LTHSC, STHSC, and proliferating states; novel p53-ROS regulatory mechanism identified | Quiescence maintenance; stem cell pool regulation | Experimental validation of predicted mechanisms | [12] |
| Epigenetic/GRN regulation | HSCs at differentiation stages | scRNA-seq, scATAC-seq | SCENIC (TF network inference); CellOracle (in silico TF perturbation) | Multi-dataset hematopoietic single-cell data | Stage-specific TF activities (GATA1, CEBPD, IRF8) validated by chromatin accessibility | TF programs governing HSC self-renewal and differentiation | Cross-validated in mouse and human hematopoiesis | [13,14] |
| Multi-omic integration | HSCs/progenitors - incomplete or unpaired multi-omics | scRNA-seq + scATAC-seq + CITE-seq | scMaui (product-of-experts VAE); totalVI (joint RNA + protein latent space) | PBMC CITE-seq benchmarks; human HSPC datasets | Robust batch correction; missing-modality handling; joint RNA + protein embedding | Systems-level HSC state characterization; HSC/LSC surface phenotype + transcriptional integration | Benchmarked on standard CITE-seq datasets; applied to HSPC/LSC contexts | [15,16] |
| B: AML and LSCs | ||||||||
| AML risk stratification | LSC-enriched compartments | Clinical, molecular, cytogenetic, and genomic features | Supervised ML (ensemble methods) | 1383 patients - multicenter cohort | Predicted complete remission and 2-year OS; outperformed ELN-only stratification | Prognosis; LSC burden surrogate; treatment-intensity guidance | External validation in an independent cohort | [17] |
| AML transcriptomic risk/epigenetic subtypes | AML blast/LSC-enriched populations | Transcriptomic + epigenetic multi-omics (multi-center) | ML epigenetic subtype classification; prognostic integration | Multi-center transcriptomic cohort | Prognostic stratification by epigenetic subtype; integration with ELN 2022 | Relapse risk; therapeutic vulnerability by epigenetic class | Multi-center validation | [18] |
| AML drug response prediction - multi-omics ensemble (MDREAM) | AML patient-derived blasts; LSC-enriched compartments | Multi-omics: Genomic mutations + transcriptomic gene expression (ex vivo drug sensitivity) | Ensemble ML (MDREAM framework; multi-omics integration) | BeatAML cohort: 278 training/183 validation; + Swedish AML cohort (n = 45); relapsed/refractory cohort (n = 12) | Spearman r = 0.68 (BeatAML validation); 77% correct responder identification at prediction confidence > 0.75 | Drug response prediction in AML patients; therapy selection; resistance identification | External multi-cohort validation (BeatAML + 2 independent cohorts) | [19] |
| AML targeted therapy - network ML | LSC populations; genomic/proteomic networks | Genomic and proteomic datasets | Network-based machine learning | Multi-dataset genomic + proteomic | Identified resistance pathways and therapeutic vulnerabilities | Immunotherapy target prioritization; resistance mechanism mapping | Not fully specified | [20] |
| AML MRD detection (MAGIC-DR) | Residual LSC-like subpopulations (immature monocytic cells) | Multiparameter flow cytometry | XGBoost + UMAP (MAGIC-DR framework) | AML MRD specimen cohort | AUC of 0.97; strong concordance with expert manual analysis; identified LSC-like residual populations missed by manual gating | Relapse risk; MRD+/- classification; sub-threshold LSC-like disease detection | Prospective validation cohort | [21] |
| AML LSC phenotyping/MRD | CD34+CD38- LSC-enriched compartment | Flow cytometry (CD117, CD34, HLA-DR) | Random forest + SHAP explainability | 194 patients (AML, AML-CR, normal BM) | Accuracy: 94.92%; AUC: 94.83%; CD117, CD34, HLA-DR top discriminators | LSC detection; early relapse identification; AML vs CR vs normal classification | Validated on a 194-patient clinical cohort | [22] |
| BCL-2/venetoclax response in LSC | LSC-enriched blasts | scRNA-seq + ex vivo drug sensitivity data | XGBoost | VenEx clinical trial patient samples | Venetoclax response prediction r = 0.71-0.84 | LSC apoptotic dependency; treatment selection for LSC eradication | VenEx clinical trial validation | [23] |
| Menin inhibition - LSC clonal architecture | KMT2A-rearranged/NPM1-mutant LSCs | Single-cell multi-omics (scRNA-seq + scATAC-seq) | DL clonal architecture mapping | Clinical trial and patient-derived samples | Mapped HOXA/MEIS1 program activity; tracked LSC self-renewal disruption under Menin inhibition | LSC self-renewal disruption; Menin inhibitor response monitoring | Clinical trial context | [24,25] |
| LSC surface target prioritization | LSC populations (CD34+) | Multiparameter flow cytometry (LSC immunophenotyping) | ML ranking/feature prioritization | Multi-dataset LSC surface marker profiling | CD123, CLL-1 identified as top immunotherapy targets vs normal HSC | Immunotherapy target selection; CAR-T/bispecific antibody design | Informed clinical trial design | [26] |
| Synergistic drug combinations - LSC | Diagnosis/relapse LSC-enriched populations | Paired scRNA-seq + ex vivo drug screens | XGBoost (combination prediction) | Paired diagnosis/relapse patient samples | Synergistic drug pair identification; clinically actionable 2-week turnaround | Personalized relapse therapy; LSC-selective combination design | Flow cytometry validation | [23] |
| AML explainable AI/precision oncology | AML patients (bulk + LSC-enriched contexts) | Transcriptomic, clinical, molecular | Explainable ML: SHAP, attention, ensemble XAI | Multi-center AML cohorts | Improved clinical interpretability; synergistic drug-response signatures via ensemble XAI | Clinical decision support; treatment selection; regulatory compliance | Multi-center validation | [27-29] |
| C: MMSCs | ||||||||
| MM spatial mapping - bone marrow niche | MMSCs; niche cells (BLIMP1+, CD8+, stromal) | Multiplex-IHC trephine biopsies | DL: MoSaicNet (tissue segmentation) + AwareNet (rare-cell detection) | MGUS and NDMM clinical trephine biopsy cohort | Spatial heterogeneity as key MGUS-to-NDMM distinction; tumor-immune spatial proximity mapped | Sanctuary identification for quiescent MMSCs; immune exclusion mapping; MGUS vs NDMM distinction | Validated on clinical MGUS and NDMM samples | [30] |
| MMSC niche interactions - treatment resistance | Quiescent MMSCs; BM stromal cells | Multi-omics datasets | Neural networks/DL | Multi-omics patient-derived datasets | Prediction of treatment resistance linked to niche-protective microenvironments | Niche-mediated drug resistance; MMSC persistence | Experimental/in vitro validation (details not fully specified) | [31,32] |
| D: Drug discovery and LSC-targeted therapy | ||||||||
| Virtual screening/AI-enhanced docking | LSC-targeted compounds; kinase/epigenetic targets | Multi-billion-compound chemical libraries (structure-based) | AI-enhanced virtual screening; molecular docking; structure-based drug design | Ultra-large compound libraries | Accelerated hit/Lead identification; 11 confirmed hits validated by X-ray crystallography; reduced screening timelines | LSC-targeted drug identification; multi-target resistance-overcoming agents | Experimental validation (X-ray crystallography of confirmed hits) | [7] |
| Drug repurposing - AML | AML/LSC targets (polypharmacological) | FDA-approved compound structures; AML molecular targets | Structure-based virtual screening; molecular dynamics simulation | FDA-approved drug library | Identified repurposable compounds with polypharmacological AML/LSC activity | Rapid clinical translation; overcoming LSC therapy resistance via repurposing | In silico validation with molecular dynamics | [33] |
| E: Biomanufacturing and cell therapy production | ||||||||
| HSC product quality control | CD34+ HSC products | Bright-field microscopy (manufacturing line) | Deep CNN (computer vision) | Clinical HSC manufacturing image dataset | High accuracy for cell viability and phenotype classification; validated vs flow cytometry | Real-time manufacturing QC; batch release decisions | Validated against the flow cytometry gold standard | [34] |
| CD34 yield prediction - cord blood | Cord blood HSCs (CD34+) | Pre/post-processing cell counts; CBU characteristics | Back-propagation ANN | 802 CBUs | 56.99% prediction accuracy for CD34 dose | Graft potency estimation; transplant planning | Validated on 802 CBUs | [35] |
| PBSC apheresis yield prediction | Autologous/allogeneic PBSC donors | Pre-apheresis CD34 counts; donor clinical characteristics | ML regression | Multi-site clinical apheresis dataset | Accurate harvest outcome prediction supporting collection scheduling | Collection efficiency; donor scheduling optimization | Clinical validation | [36] |
| Ex vivo HSC expansion optimization | HSCs in bioreactor culture | Multi-parameter sensors (O2, glucose, lactate, cytokine) | Reinforcement learning + digital twin modeling | Proof-of-concept bioreactor dataset | Optimized culture conditions; improved expansion yield via adaptive control | Expansion efficiency; GMP process optimization | Proof-of-concept validation | [37] |
| CAR-T manufacturing - smart hospital | CAR-T cell products (autologous) | Process sensor data; imaging; manufacturing records | AI + automation (Industry 4.0) | Pilot smart manufacturing hospital data | Reduced vein-to-vein time; standardized product quality | Scalable autologous CAR-T production; accessibility | Pilot implementation validation | [38] |
| Digital twin/Bioprocessing 4.0 | Cell therapy products; HSC/CAR-T | Multi-parameter bioprocess sensor streams | Digital twin simulation; reinforcement learning; AI-assisted process control | Pharma manufacturing simulation datasets | Real-time process optimization; predictive quality control | GMP compliance; scalable biomanufacturing | Simulation and proof-of-concept | [8,9,39] |
| F: HSCT - transplantation, HLA matching, and GVHD | ||||||||
| HLA genotyping extraction (NLP) | HSCT donors and recipients | Electronic health records (free text; 70000+ HLA reports) | Rule-based NLP (Python regex; 124 extraction + 8 cleaning rules) | 70000+ HLA reports (Seoul National University Hospital) | Precision: 0.892-0.999; recall: 0795-0.998 across HLA-A, B, C, DR, DQ | Donor-recipient HLA matching; adverse drug reaction prediction | Clinical validation at the SNUH registry | [40] |
| Acute GVHD prediction (CNN-NLP) | HSCT recipients | Clinical notes + HLA typing data | CNN-NLP hybrid (word2vec encoding of HLA antigens/alleles) | 18763 patients (Japanese Transplant Registry) | Stratified cumulative incidence 318% (low-risk) to 54.8% (high-risk); superior to Cox models | aGVHD risk stratification; immunosuppression planning | Registry-based validation (18763 patients) | [41] |
| Chronic GVHD phenotyping (ML + NLP) | HSCT patients with cGVHD | Clinical notes (organ involvement: Mouth, eye, liver, GI, joints, fascia, skin) | ML feature extraction + NLP clinical narrative analysis | Multicenter HSCT patient cohort | 7 distinct cGVHD phenotypes; 2.24-fold mortality difference (high vs low risk); independent of NIH severity criteria | Mortality stratification; personalized immunosuppression | Multi-center clinical validation | [42] |
| Comprehensive HSCT AI care | HSCT recipients (full care pathway) | Transplant documentation; clinical narratives; registry data | NLP + ML (composite AI review) | Multi-institutional registry and EHR data | Automated complication parsing; infection management; NRM prediction | Early warning systems; donor selection; complication management | Multi-institutional review | [4] |
Table 2 Public datasets and computational tools specifically applicable to hematopoietic stem-cell artificial intelligence research
| Stem-cell anchor | Resource name | Resource type | What it’s for (HSC/LSC-specific) | Access/Ref. |
| A: Reference atlases for HSC → lineage differentiation (where “stemness” lives) | ||||
| Human bone-marrow baseline map for projecting new normal/AML datasets and interpreting “where the cell sits” on the HSC → lineage hierarchy | BoneMarrowMap | scRNA-seq atlas + projection tool | Reference for CD34+ HSPC balance + mature compartments; projection/classification of new hematopoietic or leukemic cells | GitHub/[90] |
| Canonical mouse HSPC differentiation landscape used as benchmark for TI and lineage priming | Nestorowa et al[91] Mouse HSPC atlas | scRNA-seq | Mouse HSPC heterogeneity + lineage trajectories; common benchmark dataset | GEO GSE81682/[91] |
| Early-life/developmental hematopoiesis context | Perinatal bone marrow development atlas | scRNA-seq | Niche development + marrow colonization programs | [92,93] |
| Cord blood HSC heterogeneity (neonatal stemness programs; ex vivo culture/transplant relevance) | Human cord blood HSC populations | scRNA-seq | CD34pos vs CD34neg HSC populations; neonatal HSC state differences | GEO GSE237832/[94] |
| Human HSC activation states (quiescence → activation continuum) | Human HSC activation trajectory | scRNA-seq (SMART-seq2) | Activation trajectories in the primitive CD34+CD38-CD45RA- compartment | EGA study1 |
| Integrated RNA + chromatin programs across hematopoiesis (epigenetic stemness + TF logic) | Murine HSPC multimodal atlas | scRNA-seq + scATAC-seq | Joint transcriptomic/epigenetic differentiation programs | [95] |
| B: Multimodal phenotype ↔ transcriptome (HSC/LSC immunophenotype as a first-class signal) | ||||
| Method benchmark for RNA + protein integration (not HSC-specific, but enables HSC surface-marker integration) | PBMC CITE-seq (10X genomics) | CITE-seq (RNA + protein) | Benchmarking multimodal integration models, you then apply them to marrow HSPC/LSC | 10X datasets referenced as common benchmarks in the totalVI context[15] |
| High-parameter immune niche cell reference (HSC niche signals/immune interactions) | Murine spleen/lymph node CITE-seq | CITE-seq | Large protein panel; good demonstration dataset for multimodal modeling | [15] |
| Core computational framework for CITE-seq (HSC/LSC immunophenotyping) | totalVI | Multi-omics integration | Joint latent space (RNA + protein), batch correction, missing-protein handling → directly supports HSC/LSC surface phenotype + transcriptional state integration | [15] |
| C: Mechanistic “stemness regulators”: GRNs, TF programs, causality-leaning inference | ||||
| In silico TF perturbation to test “does this TF maintain HSC identity or drive differentiation?” | CellOracle | GRN inference + perturbation simulation | Demonstrated in mouse and human hematopoiesis; supports mechanistic framing in a stem-cell review | [14] and GitHub[96] |
| TF-activity inference from scRNA-seq to link stage-specific regulators with differentiation states | SCENIC | TF network inference | Regulatory network + regulon activity per cell; can be used to interpret HSC/LSC programs | [13] |
| D: TI (to resolve rare HSC states, branching, cycling) | ||||
| Capturing complex hematopoietic branching while preserving rare populations (useful for “rare HSC/pre-LSC” arguments) | VIA | TI | Lazy-teleporting random walks + MCMC; marketed as scalable/generalized TI | [5] |
| E: Handling missing modalities + batch effects in HSC multi omics (practical translational reality) | ||||
| Multi-omic integration when HSC datasets are incomplete/unpaired across assays | scMaui | Multi-omics integration | Product-of-experts VAE; batch + missing modality handling | [16] |
| F: Aging/activation phenotypes in HSCs (imaging + AI) | ||||
| Imaging-based “biological age” predictor from HSC nuclear architecture | ChromAgeNet | Chromatin aging prediction | CNN on 3D DAPI chromatin images; positions aging as a learnable HSC phenotype | [6] |
| G: LSC persistence proxies in patients: MRD/residual disease | ||||
| Interpretable ML that supports MRD assessment (detecting residual immature/LSC-like compartments) | MAGIC-DR | MRD detection | Interpretable ML-guided approach for AML MRD | PubMed landing[56] |
| Notable for making a large AML flow cytometry dataset publicly available for benchmarking computational MRD tools | Computational MRD assessment (GMM + novelty detection) | MRD detection | Standardization + automated MRD explicitly discusses heterogeneity issues that connect to LSC persistence framing | [97] |
Table 3 Summary of key machine learning algorithms, applications, and trade-offs in hematopoietic stem cell research
| Machine learning algorithm | Primary HSC/hematology applications | Key advantages (Pros) | Key limitations (Cons) | Ref. |
| CNN | Morphological classification of HSCs vs MPPs; 3D chromatin age prediction (ChromAgeNet); automated quality control imaging in biomanufacturing | Unparalleled performance on spatial and image data; extracts features autonomously without requiring manual gating or human-defined parameters | Highly opaque “black box” nature requiring XAI for interpretability; demands massive, accurately annotated image datasets to train | [6,11,34] |
| Random forest/ensemble trees | Flow cytometric LSC phenotyping; early relapse detection; predicting cord blood CD34+ cell yield | Highly robust to overfitting; handles tabular clinical and multi-omics data effectively; naturally provides feature importance rankings | Less effective than deep learning for highly unstructured data (like raw images or free text); can struggle with extrapolating data outside the training range | [22,35] |
| Gradient Boosting (e.g., XGBoost) | Automated MRD detection in flow cytometry (MAGIC-DR); predicting synergistic drug combinations; venetoclax response prediction | Exceptional predictive accuracy on structured clinical and omics data; handles missing data well; highly scalable for large patient cohorts | Prone to overfitting on very small sample sizes; hyperparameter tuning is complex and computationally expensive | [23,56] |
| Deep learning/ANN | Multi-omics integration (e.g., totalVI); modeling age-dependent HSC self-renewal; multi-center AML survival prediction | Can capture extremely complex, non-linear biological relationships across massive, high-dimensional datasets (e.g., integrating RNA and protein expression) | Computationally intensive; high risk of learning artifactual batch effects rather than true biology if data is not strictly harmonized | [10,15,17] |
| NLP | Extracting HLA genotypes from unstructured electronic health records; predicting and phenotyping acute and chronic GVHD from clinical notes | Unlocks vast amounts of unstructured, historical clinical data that is otherwise inaccessible to standard statistical models | Performance is heavily dependent on the quality, consistency, and language of physician documentation; it struggles with implicit clinical context | [40-42] |
| RL | Dynamic optimization of ex vivo HSC expansion; adaptive control of bioreactor parameters (cytokines, perfusion, metabolic flux) | Enables continuous, autonomous, real-time process optimization without requiring a pre-defined static protocol | Requires highly accurate “digital twins” or simulation environments to train the agent safely; initial validation in GMP environments is regulatory complex | [37,77,78] |
Table 4 Artificial intelligence performance vs conventional methods in acute myeloid leukemia and multiple myeloma risk stratification
| Disease | Conventional method | AI method | Key performance advantages | Ref. |
| AML | ELN risk stratification | Multi-omics deep learning | Outperforms ELN-based approaches. Identifies LSC burden (not quantified by ELN). > 90% accuracy in therapy resistance prediction. Integrates clinical, cytogenetic, and molecular data | [53] |
| AML | Traditional cytogenetic/molecular classification | Supervised machine learning | Superior prediction of complete remission and 2-year survival. External validation confirms generalizability | [17] |
| MM | Standard risk assessment | Multi-omics integration (neural networks) | Accuracy in therapy resistance prediction. Enhanced drug resistance prediction. Integration of genomic biomarkers and clinical parameters | [31,32,98] |
| MM | Standard histopathology | Deep learning (MoSaicNet & AwareNet) | Spatial heterogeneity detection beyond cell density. Differentiates MGUS from MM based on spatial architecture | [30] |
| HSCT GVHD | Cox proportional hazard models | CNN-NLP hybrid | Superior risk stratification. Stratifies aGVHD incidence from 31.8% to 54.8%. Processes detailed HLA information vs binary matching | [41] |
| HSCT cGVHD | NIH consensus severity criteria | Machine learning phenotyping | Identified 7 distinct phenotypes. 2.24-fold mortality difference between risk groups. Better survival stratification than traditional scores | [42] |
Table 5 Validation checklist for leukemic stem cell/hematopoietic stem cell artificial intelligence models, adapted from TRIPOD + AI and PROBAST + AI guidelines with hematopoietic stem cell/Leukemic stem cell-specific requirements
| Validation domain | Key requirement | LSC/HSC-specific considerations |
| Cohort representativeness | The training cohort must represent the target clinical population with respect to age, disease stage, and treatment era | LSC/HSC models should include balanced representation of ELN risk categories, stem-cell compartment measurements (CD34+CD38- frequencies), and both newly diagnosed and relapsed/refractory patients[52,106-108] |
| Event counts | An adequate number of outcome events (relapse, death, MRD positivity) to prevent overfitting | Minimum 10-20 events per predictor variable; for LSC-specific endpoints (e.g., LSC+ vs LSC-), ensure sufficient LSC+ cases across validation sets[52,55] |
| Internal validation | Model performance assessed on held-out data from the same source (cross-validation or hold-out split) | Report performance metrics (AUROC, calibration) separately for LSC-enriched vs LSC-depleted subgroups if the model claims to encode stemness biology[52,109,110] |
| External validation | Independent cohort from a different institution, time period, or geography | Essential for LSC models given center-to-center variability in LSC phenotyping protocols and MRD detection thresholds[52,106,110] |
| Prospective evaluation | Forward-looking validation on newly enrolled patients before clinical deployment | Required for LSC-targeted therapy selection models to confirm that AI predictions align with clinical outcomes under prospective conditions[55,109,111] |
| Dataset shift detection | Assess whether model performance degrades when applied to data with distributional differences (batch effects, assay drift) | Critical for flow cytometry-based LSC models: Validate across different antibody panels, fluorophores, and cytometers; report performance stratified by batch[55,112] |
| Calibration assessment | Predicted probabilities should match observed event frequencies | For LSC burden models, calibration plots should show agreement between predicted LSC frequency (or surrogate score) and directly measured LSC% by flow cytometry in the calibration subset[52,55,106] |
| Decision-curve analysis | Net benefit of model-guided decisions compared to treat-all or treat-none strategies | For LSC-directed therapies (venetoclax, Menin inhibitors), decision curves should quantify clinical utility across risk thresholds relevant to treatment intensification decisions[55,111] |
| Explainability and feature attribution | Use of XAI methods (SHAP, attention weights) to identify which features drive predictions | Essential for validating LSC-AI hypothesis: Determine whether high-risk predictions are driven by known stemness genes (17-gene LSC score, HOXMEIS1 programs) or alternative pathways[52,113] |
| Bias and fairness evaluation | Assess performance stratified by demographic subgroups and underrepresented populations | Evaluate whether LSC models perform equivalently across age groups (pediatric vs adult vs elderly AML), ancestry, and sex; report subgroup-specific metrics[107] |
| Missing data handling | Transparent reporting of missingness patterns and imputation strategies | LSC models often integrate multi-omics data with heterogeneous completeness (e.g., scRNA-seq available for a subset); clearly document handling of missing modalities and proteins in CITE-seq[111] |
| Comparator benchmarking | Performance compared to established clinical risk systems | For AML LSC models, benchmark against ELN 2022 risk classification, 17-gene LSC score, and LSC frequency by flow cytometry; report incremental predictive value[114] |
- Citation: Abd El Ghaffar HA, Arafat AMA, Khattab EHA, Khattab MA, Khallaf AM, Mahgoub SMA. Artificial intelligence in hematopoietic stem cell research and associated malignancies: From disease modeling to cell manufacturing. World J Stem Cells 2026; 18(8): 121077
- URL: https://www.wjgnet.com/1948-0210/full/v18/i8/121077.htm
- DOI: https://dx.doi.org/10.4252/wjsc.121077