BPG is committed to discovery and dissemination of knowledge
Minireviews
Copyright: ©Author(s) 2026.
Artif Intell Gastroenterol. Aug 8, 2026; 7(2): 118476
Published online Aug 8, 2026. doi: 10.35712/aig.118476
Table 1 Cross-sectional imaging for evaluation of biliary strictures
Imaging
Sensitivity range/specificity range, %
Notes
CT scan[15-17]Sensitivity 75-80. Specificity 60-80Study context: Prospective, blinded comparisons in patients with obstructive jaundice. Patient population: Mixed benign & malignant strictures (distal and hilar). Reference standard was histology or > 12 months[16,17]. Primary role: Excellent for staging, vascular assessment, and detecting metastases. Key limitation: Limited specificity for differentiating malignant from benign strictures, particularly in the absence of a discrete mass
MRCP[18]Sensitivity 83-90. Specificity 94-98Study context: Prospective observational study comparing MRCP directly to ERCP as a reference standard. Patient population: 60 patients with suspected CBD or pancreatic duct pathologies. Primary role: Non-invasive gold standard for evaluating biliary tree anatomy, stricture morphology, and level of obstruction. High accuracy for choledocholithiasis and ductal dilation. Key limitation: Provides anatomical, not histologic, diagnosis
18F-fluorodeoxyglucose positron emission tomography[19]Sensitivity 85-92. Specificity 51-90Study context: Pooled data from a meta-analysis. Patient population: 47 studies (n = 2125 patients). Primary role: Not for primary stricture characterization. High utility for staging (lymph node/distant metastasis). Changes management in approximately 15% of cases, primarily via upstaging. Key limitation: Very low specificity (51%) for diagnosing malignancy at the primary site due to false positives from inflammation (e.g., PSC, cholangitis). Cannot replace histopathological confirmation
Table 2 Endoscopic retrograde cholangiopancreatography-based sampling modalities
Baseline modality
Sensitivity and specificity range, %
Combined (with biopsy) sensitivity and specificity range, %
Incremental gain, %
Notes
Brush cytology alone[1,11,13,20-22]Sensitivity 27-79. Specificity 90-100Sensitivity 66-83.3. Specificity 95-100Incremental 15-20 sensitivity over single methodsData from retrospective cohort (Nanda et al[13] n = 61, CCA diagnosis including hilar, perihilar and distal strictures) and RCT comparing brush types (n = 64, extrahepatic strictures[21]). Data from systematic review and meta-analysis (16 studies); yield improves with multiple passes and bile cytology, RCT shows yield improves with modified biopsy forceps[20,22]
Brush cytology and FISH][1,13,23-26]Sensitivity 35-84. Specificity 54.1-98Sensitivity 70-82. Specificity 90-10020-40 yield gain over cytology alone, highest ERCP yield (triple sampling), real-time risk stratification Data from retrospective cohort (Nanda et al[13] n = 61, CCA diagnosis including hilar, perihilar, and distal strictures), comparative study by Kipp et al[23]. n = 131, including biliary strictures, data from prospective study, n = 81 with biliary and pancreatic duct strictures[24], data from retrospective study n = 614 patients[25], 10 years retrospective study of 281 patients[26]. FISH detects polysomy (aneuploidy) in 80% malignancies, misses diploid tumors, biopsy has no submucosal access
Table 3 Advanced modalities for biliary strictures
Modality
Sensitivity range, %
Specificity range, %
Advantages/limitations
Notes
Intraductal ultrasound[33-35]93-9779-90Pros: Real-time wall-layer/T-staging, superior for proximal strictures. Cons: Probe insertion issues, limited N-staging, advanced tumors Large retrospective cohorts (n = 193; n = 379) using histopathology or long-term follow-up as reference standards[34,35]
Cholangioscopy directed biospy[28,36,37] 58-8690-100Pros: Direct visualization, targeted sampling in proximal lesions. Cons: Costly equipment, expertise required, distal limitations Multicenter retrospective SpyGlass DS™ cohort (n = 206)[36]; pooled evidence from systematic review and meta-analysis extracted data from 15 studies (n = 539) by Badshah et al[37]
EUS + FNA[38-40]76-9490-97Pros: Excellent for distal strictures and mass lesions, deeper access. Cons: Perihilar challenges, lower yield without mass. Significant impact on surgical decision-makingMeta-analysis (6 studies, n = 497)[38] and prospective cohorts (n = 50)[39]; n = 44[40] confirm high diagnostic accuracy of EUS-FNA, particularly for extraductal lesions > 1.5 cm. The gold standard was surgery or 6 months follow up
EUS guided FNA + ERCP guided TA[38,39]86-9898-100Pros: Maximizes yield in the same session, reduces false negatives. Cons: Procedural time and risk increase, resource-intensive Systematic review and meta-analysis of same-session procedures (6 studies, n = 497) demonstrated an accuracy of 96.5%, superior to either modality alone, including hilar, perihilar, and distal strictures[38]. Prospective comparative study (n = 50) showed combined sensitivity 97.9% and accuracy 98%, significantly reducing false negatives[39]. The gold standard was surgery or 6 months follow up
EUS + FNB[41,42]85-9890-100Pros: Better tissue yield than FNA, ideal for distal strictures. Cons: Seeding risk, operator-dependentProspective multicenter cohort (n = 465) showed superior tissue core yield and histologic accuracy with 22G FNB (99%) vs FNA (61%), with reduced blood contamination; combined analysis reached 100% diagnostic accuracy[41]
EUS guided FNB and ERCP guided tissue acquisition[43,44,60]83-10095-100Pros: EUS-FNB serves as a first-line or complementary tool for diagnosing indeterminate biliary strictures, offering superior sensitivity (83%-98%) over ERCP sampling, especially for distal/extrahepatic lesions without visible masses. Cons: Increased procedural time/risk, dual expertise neededRetrospective BS cohort (n = 51) with surgical histology or radiologic/clinical follow-up as gold standard showed higher accuracy with same-session EUS-FNB + ERCP (83.3% sensitivity; 87.5% accuracy; 100% specificity) vs either alone[43]. Sensitivity drops to 56%-75% with stents or hilar location due to access challenges[44]
pCLE[45] 75-8080-85Pros: In vivo histology, real-time neoplasia detection, improved with Paris Classification. Cons: Probe fragility, steep learning curve, limited availability, high costValidation study of the Paris Classification using 40 pCLE sequences (19 inflammatory, 6 benign, 15 malignant indeterminate biliary strictures) demonstrated improved specificity with maintained overall accuracy (approximately 82%). The gold standard was histopathology or clinical follow-up. Interobserver agreement was fair (κ = 0.37). Lesions involved indeterminate bile duct strictures, primarily the extrahepatic biliary tree during ERCP-based evaluation
OCT/VLE[46,47] 7969Pros: Microstructural imaging of desmoplasia excels in tight strictures. Cons: Interpretive variability, limited availability, large RCTs needed to confirm its clinical impactProspective ERCP-OCT study (n = 37 biliary strictures; 19 malignant, 16 benign) using histology/EUS-FNA/surgery or ≥ 12-months follow-up as reference standard showed sensitivity 79% and specificity 69% (≥ 1 criterion); specificity 100% when both criteria were required. Combined with brushings, increased sensitivity to 84%[47]
Table 4 Artificial intelligence based radiomics models for non-invasive diagnosis and risk stratification of biliary strictures
Study year
AI model
Data source
Key performance
Clinical impact
2025, meta nalysis[56] Various ML, LR, RF, SVM, DLCT (13 studies), EUS (5 studies), MRI (3 studies), PET-CT (3 studies), 24 case-control studies; 14406 patients (7635 PDAC)Sensitivity 0.92 (95%CI: 0.91-0.94); Specificity 0.90 (95%CI: 0.85-0.94); AUC 0.94 (95%CI: 0.74-0.99); DOR 110 (95%CI: 62-194)Non-invasive PDAC screening; reduces EUS-FNA need; CT best for initial triage, limited by moderate radiomics quality score and case-control design
Multi-institutional radiomics study[52] ML classifiersCT imaging of PDACs < 2 cmSensitivity > 90% for small PDAC detection; cross-population generalizability demonstratedFirst to characterize PDAC-specific radiomic signature (decreased intensity, increased NGTDM busyness) linking imaging features to desmoplastic heterogeneity; validated across multiple populations
Differentiation AIP vs PDAC[53]. Radiomics-based MLThin-slice venous-phase CTAccuracy 95.2%; AUC 0.975; sensitivity 89.7%; specificity 100%Demonstrated superiority of thin-slice venous-phase imaging; proved AI can discriminate malignancy from complex benign inflammatory conditions (previously a major limitation)
PET/CT radiogenomics[54,55] Radiomic analysisFDG PET/CT with genomic correlationMetabolic texture features correlate with KRAS and SMAD4 mutationsFirst demonstration that radiomic phenotypes reflect underlying tumor genotype; provides non-invasive window into tumor biology beyond structural imaging
2023, narrative review[57]Radiomics review (various ML)CT/MRI across HPB studiesPDAC detection/differentiation: AUC 0.71-0.99. Tumor grading prediction: AUC 0.73-0.90. IPMN high-grade dysplasia prediction: AUC up to 0.84Supports early PDAC screening, cyst risk stratification (e.g., IPMN high-grade dysplasia AUC 0.84), and resectability assessment in pancreatic/HCC/ICC
Table 5 Limitations of artificial intelligence and potential mitigation strategies in the diagnosis of biliary strictures
Limitations
Evidence from current literature
Illustration from AI biliary stricture studies
Potential mitigation strategies
Retrospective design and selection bias[14,74] The majority of published AI studies are retrospective and single-center, limiting generalizability and inflating performance estimatesZhang et al[71] (MBSDeiT) demonstrated prospective real-time D-SOC prediction on a retrospectively trained model, but did not assess the impact on clinical outcomes or compare performance with biopsy guidance. Saraiva et al[68] similarly acknowledged that despite their large dataset for a proof-of-concept study, clinical validation requires a much larger volume of dataConduct prospective, multicenter randomized trials and pragmatic validation studies
Black-box decision making[12,68] Most high-performing CNN and ensemble models lack intrinsic interpretability, reducing clinician trust in AI-guided lesion targeting and biopsy decisionsSaraiva et al[68] demonstrated post-hoc heatmap validation of CNN attention to tumor vessels and papillary projections but lacked real-time interpretability during procedures, similarly acknowledged that despite their large dataset for a proof-of-concept study, clinical validation requires a much larger volume of dataIntegrate explainable AI techniques (e.g., Grad-CAM, SHAP) as real-time overlays on D-SOC/EUS. Develop AI systems that highlight and verbalize decision-driving features
Limited dataset size and diversity[70,74,88] Stricture-specific EUS and cholangioscopy datasets remain small (< 10000 labelled images across studies), with underrepresentation of perihilar and benign inflammatory stricturesRobles-Medranda et al[70] conducted a two-phase multicenter study (n = 164), yet disease spectrum and procedural variability remained limited, with ongoing discrepancy between operators visual impression using current classifications for indeterminate biliary stricturesCreate open-access, multi-vendor endoscopic image repositories (e.g., expanding The Cancer Imaging Archive. Employ federated learning to train models across institutions without sharing raw data
Domain shift and device dependency[70,89] Model performance may degrade when applied across different EUS, D-SOC processors, probes, contrast agents, or imaging protocolsThe multicenter validation by Robles-Medranda et al[70] may be affected by variability in D-SOC systems and protocolsImplement cross-platform training, harmonization of acquisition protocols, and external validation across vendors and geographic regions
Lack of clinical workflow integration[86,87] Many AI models are evaluated offline on static images or pre-recorded videos, without assessment of real-time feasibility, procedural impact, or endoscopist interactionOnly Marya et al[87] implemented and validated real-time D-SOC classification during live proceduresConduct human-factors studies on endoscopist-AI interaction, and conduct real-time deployment trials measuring procedural metrics
Absence of outcome-driven validation[84,85] Most studies report diagnostic accuracy but lack data on downstream clinical outcomes (time to diagnosis, avoided ERCP, surgical yield)A multicenter AI study[85] demonstrated superior diagnostic accuracy but did not report the impact on avoided ERCPs or surgical outcomesDesign endpoint-driven trials linking AI-guided diagnosis to clinical outcomes. Perform formal cost-effectiveness and patient-centered measures


Write to the Help Desk