BPG is committed to discovery and dissemination of knowledge
Minireviews
Copyright: ©Author(s) 2026.
Artif Intell Gastroenterol. Aug 8, 2026; 7(2): 116570
Published online Aug 8, 2026. doi: 10.35712/aig.v7.i2.116570
Table 1 Synthesis of literature data and key artificial intelligence advancements in celiac disease histopathological diagnosis
AI application
Primary AI methodology
Key impact on CD histopathological diagnosis
Performance metrics1
Ref.
Challenges/future directions
Automated VCR measurementSemantic Segmentation (U-Net, FCN), CNNsObjective, reproducible quantification of villous atrophy and crypt hyperplasia; reduces inter-observer variability and manual effort.Dice coefficient: 0.87-0.93; VCR Agreement with experts: 0.91-0.97; mean absolute error in VCR: 0.08-0.15[39,41]Robustness to diverse staining/scanning protocols; fine-tuning for subtle partial atrophy; validation across multicenter datasets
IEL quantificationObject detection (YOLO, SSD), instance segmentation (mask R-CNN)Accurate, rapid counting and density estimation of IELs; aids in Marsh 1 diagnosis and differentiation from other lymphocytosis causes.YOLO mAP: 85%-92%; SSD mAP: 80%-88%; mask R-CNN instance accuracy: 88%-94%; counting error reduction: 65%-82% vs manual enumeration[34,36]Differentiation of IELs from other inflammatory cells; consistency across varying tissue preparations; clinical correlation with specific IEL thresholds and establishment of standardized reference ranges
Automated marsh classificationCNNs (ResNet, Inception), MILEnables automated or semi-automated grading of mucosal damage (Marsh 0-3c); high sensitivity/specificity comparable to expert pathologistsOverall diagnostic accuracy: 89%-96%; sensitivity: 85%-94%; specificity: 87%-98%; AUC-ROC: 0.92-0.98; agreement with expert consensus (κ value): 0.85-0.94[23,45]Interpretability of “black-box” models (XAI); generalizability across unseen datasets/ethnicities; handling of equivocal or patchy lesions; ethical considerations regarding AI in clinical decision-making
Early disease detection/subtlety enhancementAttention mechanisms in CNNs, anomaly detectionHighlights subtle changes in architecture or IELs in early CD; acts as a “second reader” to reduce missed diagnosesSensitivity for Marsh 1 detection: 82%-91%; sensitivity for early Marsh 3a: 79%-88%; Reduced missed diagnosis rate: 35%-48% vs single-reader evaluation[20,53]Validation against long-term patient outcomes for early detection; integration into pathologist workflow for seamless alert generation; prospective clinical trials
Digital pathology integration and workflowCloud-based AI platforms, API integrationStreamlines workflow from WSI acquisition to AI analysis and reporting; enables scalability and remote access.Processing time reduction: 65%-80% vs manual review; WSI analysis throughput: 20-40 slides/hour; user satisfaction scores: 7.8-8.9/10[18,54]Interoperability with existing systems; data privacy and security compliance (General Data Protection Regulation, Health Insurance Portability and Accountability Act); infrastructure requirements for WSI storage and processing; cost-effectiveness analyses
XAIGrad-CAM, LIME, attention mapsIncreases transparency of AI decisions; builds trust and aids pathologists in understanding model rationale (e.g., heatmaps)Pathologist concordance with XAI explanations: 81%-89%; Trust scores before vs. after XAI implementation: 5.2 → 7.8/10; Time to interpret AI output: Reduced by 40%-55%[55,56]Developing clinically meaningful explanations; standardization of XAI outputs for diagnostic review; ensuring XAI insights are actionable for pathologists; balancing complexity and interpretability: “Pathologist-in-the-loop” validation
Data augmentation/synthetic data generationGANs, traditional augmentation techniquesExpands training datasets, improves model robustness and generalization, especially for rare findings or limited dataDataset expansion: 4-8 fold with traditional augmentation; Model accuracy improvement: 6%-18% with GAN-generated data (median 11%); expert validation of synthetic images: 71%-89%[49,51]Ensuring pathological fidelity of synthetic data; preventing generation of misleading artifacts; ethical implications of using synthetic data in diagnostic training; regulatory framework development
Weakly supervised learningMILReduces annotation burden by leveraging slide-level labels for training; useful for large datasets where fine-grained annotation is impracticalSlide-level diagnostic accuracy: 86%-92%; annotation time reduction: 70%-85%; instance-level precision: 79%-87%; AUC-ROC: 0.88-0.96[43,45]Precision in localizing specific pathological features from weak labels; risk of model focusing on non-diagnostic features; validation of learned attention patterns; threshold establishment for clinical deployment


Write to the Help Desk