Published online Sep 7, 2026. doi: 10.3748/wjg.119646
Revised: April 9, 2026
Accepted: April 28, 2026
Published online: September 7, 2026
Processing time: 189 Days and 21.8 Hours
Colorectal cancer largely arises from advanced adenomas (AA), yet current noni
To develop and validate a tongue image-clinical nomogram for AA detection and compare it with conventional risk scores.
This prospective observational study randomly divided 880 adults into training and validation cohorts (7:3). Quantitative tongue phenotypes were extracted via deep learning from standardized images. Univariate analysis, least absolute shrinkage and selection operator regression, and multivariate analysis established the nomogram. Calibration curves, the area under the receiver operating characteristic curve (AUC), and decision curve analysis assessed discrimination, accuracy, and clinical utility. Additionally, fixed-specificity classification was applied within the high-risk (APCS score ≥ 4) subgroup.
In the training cohort (n = 616; 226 AA), the nomogram achieved an AUC of 0.739, significantly superior to the tongue-only model (0.694), APCS score (0.663), and modified APCS (M-APCS) score (0.658) (all P < 0.002). In the validation cohort (n = 264; 92 AA), the nomogram maintained an AUC of 0.711, exceeding conventional scores (AUC: 0.640-0.654). Within the high-risk subgroup (APCS score ≥ 4; n = 213), the nomogram retained moderate discrimination (AUC: 0.75-0.76) and performance comparable to the tongue-only model, whereas conventional scores approached chance levels (AUC: 0.52-0.57). Using cutoffs targeting approximately 80% specificity in the high-risk validation subset, the nomogram detected 58.33% of AA vs 19.44% for APCS score and 25.00% for M-APCS score.
The tongue image-clinical nomogram may complement conventional scores for noninvasive AA risk refinement, especially in APCS-defined high-risk individuals.
Core Tip: This study developed and internally validated a noninvasive nomogram integrating quantitative tongue phenotypes with clinical factors for advanced adenoma detection. In the validation cohort, the nomogram showed higher discrimination than the Asia-Pacific colorectal screening (APCS) and modified APCS scores. In the high-risk subgroup (APCS ≥ 4), where conventional scores had limited discriminative value, the nomogram maintained moderate performance and detected 2-3 times as many advanced adenomas as the conventional scores did at approximately 80% specificity, suggesting potential as a noninvasive risk-refinement tool for opportunistic screening.
- Citation: Zhang J, Zhu ML, Xie YX, Fan YQ, Zhang W, Liu SS, Yu XM, Wang W, Fu XY. Development and validation of a tongue image-clinical nomogram for detection of advanced colorectal adenoma. World J Gastroenterol 2026; 32(33): 119646
- URL: https://www.wjgnet.com/1007-9327/full/v32/i33/119646.htm
- DOI: https://dx.doi.org/10.3748/wjg.119646
Colorectal cancer (CRC) remains a major cause of cancer morbidity and mortality worldwide[1]. Many CRCs arise through a multistep adenoma-carcinoma sequence, and advanced adenoma (AA) is an important precursor lesion for secondary prevention[2]. Epidemiologically, AAs are detected in approximately 5%-10% of average-risk individuals undergoing screening and harbor a substantially higher risk of malignant transformation compared to nonadvanced adenomas (NAA)[3]. Because AAs are often asymptomatic, their identification depends on screening rather than symp
Tongue diagnosis has long been used in traditional Chinese medicine (TCM) to describe visible oral manifestations associated with internal conditions[10,11]. From a modern biomedical perspective, several observations provide a ra
A major limitation of conventional tongue inspection is subjectivity. Computerized tongue image analysis offers a more standardized approach by converting visual features into quantitative descriptors of color and texture[21,22]. Recent studies have reported associations between tongue-image markers and established cancers, including gastric cancer[23,24]. However, whether such digital phenotypes can help identify premalignant AA, and whether they add predictive value beyond established clinical risk scores, remains unclear.
We conducted a prospective observational study to develop and internally validate a nomogram integrating quanti
This single-center prospective observational study was conducted at the Second People’s Hospital Affiliated with Fujian University of Traditional Chinese Medicine (Fujian Province, China) from March to December 2025, and was reported in accordance with the STROBE statement. During recruitment, adults aged 18-70 years who were scheduled for colono
Tongue images were collected before colonoscopy using a standardized tongue-imaging device (Portable Intelligent TCM Inspection Instrument; MQ-SXZN-ZA1) under a uniform protocol (photographs demonstrating the device and the standardized patient acquisition process are provided in Supplementary Figure 1). To eliminate inter-operator variability and ensure consistent imaging positioning, all image acquisitions were performed by a single trained clinical investigator throughout the study. To minimize short-term behavioral effects on tongue coating and color, participants were in
Using the above pipeline, we extracted 157 quantitative tongue-image features as candidate predictors, computed from the whole-tongue region as well as from the automatically separated and segmented tongue body and tongue coating regions. To enable region-specific quantification, features were computed for the whole tongue and five predefined subregions: Anterior (tip), central, posterior (root), and bilateral margins. Overall, the feature set covered two major domains: (1) Color descriptors summarizing pixel-intensity distributions in standardized color representations; and (2) Texture/surface-pattern descriptors capturing coating thickness-related patterns and local spatial variation. Color features were computed in hue, saturation, value (HSV) and International Commission on Illumination (CIELAB) color spaces; in HSV, V denoted brightness (value) and H hue (color tone), whereas in CIELAB, a* and b* represented the red-green and yellow-blue axes, respectively. The reported features specify the corresponding image channel alongside the anatomical region (e.g., “central tongue-body brightness” corresponds to the extraction from the HSV V channel). The coating-thickness feature was an image-derived descriptor reflecting the degree of thick tongue coating rather than a direct physical measurement.
Throughout the manuscript, tongue-image features are reported using standardized, descriptive labels; the correspondence between code variable names and the labels used in the text and figures is provided in Supplementary Table 1 (variable dictionary) to ensure reproducibility. To ensure clinically meaningful interpretability and avoid extreme odds ratios (ORs) associated with full-range transitions (i.e., from 0 to 1), continuous predictors naturally bounded within the[0,1] interval were rescaled to a percentage scale (multiplied by 100). Consequently, the reported ORs reflect the risk change associated with a 1-percentage-point increase (a 0.01 increment on the original scale), whereas predictors outside the 0-1 range were maintained on their original scales. This rescaling strategy modifies only the unit of interpretation without affecting model fit, discrimination, calibration, or statistical inference.
For each participant, we calculated the APCS and M-APCS scores using routinely collected variables according to published algorithms, and used them as conventional comparators. The APCS score assigned points for age (< 50 years, 0; 50-69 years, 2; ≥ 70 years, 3); sex (male, 1; female, 0); first-degree family history of CRC (yes, 2; no, 0); and smoking (ever/current, 1; never, 0). Total scores ranged from 0 to 7. Risk strata were defined as low (0-1), intermediate (2-3), and high (4-7)[7]. The M-APCS uses age (50-54 years, 0; 55-64 years, 1; 65-70 years, 2); sex (male, 1; female, 0); first-degree family history of CRC (yes, 1; no, 0); smoking (yes, 1; no, 0); and body mass index (BMI) (< 23 kg/m2, 0; ≥ 23 kg/m2, 1). This yielded total scores from 0 to 6. Risk strata were defined as low (0), medium (1-2), and high (3-6); M-APCS score ≥ 3 was considered high risk[8]. Both scores were treated as fixed rule-based predictors and compared with tongue-based models for discrimination and clinical utility, as prespecified.
Models were developed in the training cohort and evaluated in an internal validation cohort, with AA as the outcome. Sample size was planned a priori using the events-per-variable (EPV) rule for multivariable logistic regression (EPV ≥ 10) applied to the training cohort, where model coefficients were estimated. Assuming k predictors in the final multivariable model, we targeted at least 10 × k AA events in the training set to support stable model estimation. First, tongue-image features were screened in the training cohort to identify candidates associated with AA. Univariate analysis was per
We performed a prespecified subgroup analysis in a high-risk screening population defined as APCS score ≥ 4. Parti
For fair threshold-dependent between-model comparisons within the APCS score ≥ 4 subgroup, we prespecified a fixed-specificity thresholding strategy. After extracting APCS score ≥ 4 participants within each set, we determined one operating cutoff per method using only the APCS score ≥ 4 training subset, targeting approximately 80% specificity and thereby approximating a common false-positive burden across methods. For probability-based models, the cutoff was chosen on the predicted-risk scale to achieve training-set specificity closest to 80%. For score-based methods, cutoffs were chosen on the integer score scale; because only discrete thresholds are available, we selected the value yielding specificity closest to 80% in the training subset. Cutoffs were not reoptimized in validation; each training-derived cutoff was applied unchanged to the APCS score ≥ 4 validation subset to avoid optimistic bias. Classification performance was summarized using confusion-matrix counts (true positive, false positive, true negative and false negative), with AA as positive and NAA as negative. Using these counts, we calculated sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and accuracy for each method in both cohorts to provide clinically interpretable, threshold-dependent comparisons under a matched false-positive rate.
The protocol was approved by the Ethics Committee of the Second People’s Hospital Affiliated with Fujian University of Traditional Chinese Medicine (No. SPHFJP-Y2025106-01). Written informed consent was obtained from all participants.
All statistical analyses and data visualizations were performed using R (version 4.4.2). Continuous variables are presented as median (interquartile range) and compared using the Mann-Whitney U test. Categorical variables are presented as n (%) and compared using the χ2 test or Fisher’s exact test. Univariate analysis and multivariable logistic regression analyses were used to identify factors associated with AA. For feature reduction, LASSO logistic regression was implemented using the glmnet package with cross-validation, applying the one-standard-error criterion (λ 1se) to select a parsimonious predictor set. Nomograms and calibration curves were developed using the rms package, and nomograms were visualized using regplot. Discrimination was evaluated using ROC curves and AUCs, and correlated AUCs were compared using DeLong’s test (pROC package). Clinical utility was assessed using DCA (rmda package). Bootstrap resampling (500 repetitions) was used for calibration correction and to assess the stability of net-benefit estimates. All tests were two-sided, and P < 0.05 was considered statistically significant.
After eligibility screening, 880 participants were included (Figure 1). Participants were randomly split (7:3) to a training cohort (n = 616) and a validation cohort (n = 264). The training cohort included 390 participants with NAA and 226 with AA; the validation cohort included 172 with NAA and 92 with AA.
Baseline characteristics are summarized in Table 1. In the training cohort, compared with the NAA group, participants with AA were more likely to be male (61.9% vs 42.3%) (P < 0.001), older [58.0 (50.0-67.0) years vs 54.0 (42.0-62.8) years] (P < 0.001); and ever smokers (former/current) (35.0% vs 15.6%) (P < 0.001). A family history of CRC was also more common in the AA group (12.4% vs 6.7%) (P = 0.023). APCS and M-APCS scores were also higher in the AA group than in the NAA group (both P < 0.001), whereas BMI did not differ significantly (P = 0.065). Similar patterns were observed in the validation cohort. Participants with AA were more likely to be male (60.9% vs 43.6%) (P = 0.011); older [60.0 (51.8-66.2) years vs 54.5 (44.0-62.0) years] (P = 0.001); and ever smokers (34.8% vs 20.9%) (P = 0.021). APCS and M-APCS scores remained higher in the AA group (both P < 0.001), while BMI (P = 0.351) and family history of CRC (P = 0.879) were not significantly different between AA and NAA. Baseline characteristics were generally comparable between the training and validation cohorts (Table 1), supporting the appropriateness of using the training cohort for model development and the validation cohort for performance assessment.
| Variable | Training cohort (n = 616) | Validation cohort (n = 264) | 1P value | ||||
| NAA (n = 390) | AA (n = 226) | P value | NAA (n = 172) | AA (n = 92) | P value | ||
| Sex | < 0.001 | 0.011 | 1.000 | ||||
| Female | 225 (57.7) | 86 (38.1) | 97 (56.4) | 36 (39.1) | |||
| Male | 165 (42.3) | 140 (61.9) | 75 (43.6) | 56 (60.9) | |||
| Age (year) | 54.0 (42.0-62.8) | 58.0 (50.0-67.0) | < 0.001 | 54.5 (44.0-62.0) | 60.0 (51.8-66.2) | 0.001 | 0.796 |
| BMI (kg/m2) | 23.5 (21.2-25.6) | 24.0 (22.3-25.7) | 0.065 | 23.7 (21.2-26.3) | 24.1 (22.0-26.0) | 0.351 | 0.494 |
| Smoking status | < 0.001 | 0.021 | 0.377 | ||||
| Never | 329 (84.4) | 147 (65.0) | 136 (79.1) | 60 (65.2) | |||
| Ever (former/current) | 61 (15.6) | 79 (35.0) | 36 (20.9) | 32 (34.8) | |||
| Family history of colorectal cancer | 0.023 | 0.879 | 0.155 | ||||
| No | 364 (93.3) | 198 (87.6) | 163 (94.8) | 86 (93.5) | |||
| Yes | 26 (6.7) | 28 (12.4) | 9 (5.2) | 6 (6.5) | |||
| APCS score | 2.0 (1.0-3.0) | 3.0 (2.0-4.0) | < 0.001 | 2.0 (1.0-3.0) | 3.0 (2.0-4.0) | < 0.001 | 0.896 |
| M-APCS score | 2.0 (1.0-3.0) | 3.0 (2.0-4.0) | < 0.001 | 2.0 (1.0-3.0) | 3.0 (2.0-4.0) | < 0.001 | 0.569 |
To construct a parsimonious tongue-image-based model for identifying AA, we evaluated all 157 tongue-image features using univariate analysis in the training cohort. With a prespecified threshold of P < 0.05, 52 features were retained as candidates for penalized regression (Supplementary Table 2). Given the high dimensionality and potential collinearity among tongue-image variables, these candidates were further subjected to LASSO logistic regression. In the coefficient path plot, increasing penalization (larger λ) progressively shrank most coefficients toward zero (Figure 2A). The optimal penalty parameter was determined by cross-validation using the binomial deviance curve (Figure 2B), yielding λmin = 0.0122661 and λ 1se = 0.04951844. To favor a more parsimonious model with improved potential generalizability, λ 1se was selected, resulting in seven features with non-zero coefficients: Tongue-coating thickness; Central tongue redness (CIELAB a* channel); Anterior tongue brightness (HSV V channel); Central tongue-body brightness (HSV V channel); Posterior tongue-body hue (HSV H channel); Posterior tongue-coating yellowness (CIELAB b* channel); And anterior tongue-coating redness (CIELAB a* channel) (Figure 2). These LASSO-selected features were simultaneously entered into a multivariable logistic regression model (enter method) to identify predictors independently associated with AA, aiming to identify independently associated tongue-image features with AA after mutual adjustment and to obtain interpretable effect estimates (ORs and 95%CIs). Based on the multivariable results, three nonsignificant features were removed (all P > 0.05), and the final tongue-only model retained four independent tongue-image predictors (Figure 3A): Tongue-coating thickness (OR = 2.887, 95%CI: 1.823-4.602) (P < 0.001); Anterior tongue brightness (OR = 0.93, 95%CI: 0.868-0.991) (P = 0.029); Central tongue-body brightness (%) (OR = 0.946, 95%CI: 0.907-0.987) (P = 0.011); And posterior tongue-coating yellowness (OR = 1.127, 95%CI: 1.034-1.231) (P = 0.007). These predictors were carried forward for subsequent model development and comparative evaluation.
Building on the tongue-only predictors, we developed a tongue-clinical model by integrating selected tongue-image features with prespecified clinical risk factors (age, sex, smoking status, and family history of CRC) in the training cohort. These clinical variables were entered a priori because they are routinely available determinants used in screening practice and constitute core components of conventional risk scores (APCS and M-APCS). Prespecification improves interpreta
All candidate predictors were entered simultaneously into a multivariable logistic regression model (enter method). Tongue-coating thickness (OR = 3.213, 95%CI: 2.162-4.801) (P < 0.001); central tongue-body brightness (%) (OR = 0.954, 95%CI: 0.913-0.998) (P = 0.038); and posterior tongue-coating yellowness (OR = 1.158, 95%CI: 1.062-1.265) (P = 0.001) remained significant tongue-image predictors of AA (Figure 3B). Among clinical predictors, age was associated with increased odds of AA per 1-year increment (OR = 1.021, 95%CI: 1.006-1.036) (P = 0.008); ever smoking was associated with AA (OR = 2.106, 95%CI: 1.299-3.428) (P = 0.003); and family history of CRC was also significant (OR = 1.939, 95%CI: 1.050-3.585) (P = 0.034). Male sex showed a borderline association (OR = 1.496, 95%CI: 0.978-2.290) (P = 0.063) but was retained given its established role as a demographic risk factor in colorectal neoplasia. Anterior tongue brightness showed borderline significance (P = 0.059) and was not retained in the final tongue image–clinical model. Given its lack of significant between-group differences in our cohort, and because its inclusion in the M-APCS scoring system is debated, BMI was omitted from the multivariable model to maintain strict parsimony and avoid introducing statistical noise. M-APCS score was still calculated strictly according to its original definition (including BMI) for comparative evaluation. Accordingly, the final tongue-clinical model comprised seven predictors: Age, sex, smoking status, family history of CRC, tongue-coating thickness, central tongue-body brightness, and posterior tongue-coating yellowness (Figures 3 and 4). To facilitate clinical application, we constructed a nomogram based on the final model (Figure 4). Each predictor was assigned a point value, and the points were summed to generate a total score corresponding to an individualized pre
Performance of the tongue image-clinical nomogram: We assessed the performance of the tongue image-clinical nomogram in the training and validation cohorts (Figure 5). ROC analysis showed moderate discrimination, with an AUC of 0.739 (95%CI: 0.697-0.780) in the training cohort (Figure 5A) and AUC of 0.711 (95%CI: 0.645-0.778) in the vali
Comparison with the tongue-only model and conventional risk scores: We compared the discriminative performance of the tongue image-clinical nomogram, tongue-only model, APCS score, and M-APCS score in both cohorts (Figure 6A and B). In the training cohort, the nomogram showed the highest discrimination (AUC = 0.739, 95%CI: 0.697-0.780), outperforming the tongue-only model (AUC = 0.694, 95%CI: 0.649-0.738), APCS score (AUC = 0.663, 95%CI: 0.620-0.706), and M-APCS score (AUC = 0.658, 95%CI: 0.614-0.701) (Figure 6A). Pairwise comparisons using DeLong’s test for correlated ROC curves confirmed that the nomogram achieved significantly higher AUC than the tongue-only model (Z = 3.122, P = 0.001794), APCS score (Z = 3.868, P = 0.0001098), and M-APCS score (Z = 3.806, P = 0.0001415).
In the validation cohort, the nomogram maintained the best discriminative ability (AUC = 0.711, 95%CI: 0.645-0.778), remaining numerically higher than the tongue-only model (AUC = 0.676, 95%CI: 0.605-0.747), APCS score (AUC = 0.654, 95%CI: 0.588-0.720), and M-APCS score (AUC = 0.640, 95%CI: 0.573-0.707) (Figure 6B). However, DeLong’s test did not detect a significant difference between the nomogram and the tongue-only model in the validation cohort (Z = 1.432, P = 0.1521), while the nomogram showed borderline superiority over APCS score (Z = 1.950, P = 0.05124) and significantly higher AUC than M-APCS score (Z = 2.453, P = 0.01419). Overall, discrimination decreased modestly from training to validation across all approaches, but the tongue image–clinical nomogram consistently ranked first.
DCA demonstrated higher net benefit for the nomogram than for APCS and M-APCS across a broad range of threshold probabilities in both cohorts, with comparable or greater net benefit than the tongue-only model over clinically relevant thresholds (Figure 6C and D). Calibration curves for the four approaches are shown in Figure 6E and F and indicated acceptable agreement between predicted and observed risks overall. To quantify incremental risk classification, NRI and IDI were calculated for the nomogram vs the conventional scores. Compared with APCS, the nomogram improved reclassification in the training cohort (categorical NRI = 0.124, 95%CI: 0.049-0.200, P = 0.001; continuous NRI = 0.565, P < 0.001; IDI = 0.100, 95%CI: 0.074-0.126, P < 0.001) and in the validation cohort (categorical NRI = 0.298, 95%CI: 0.181-0.416, P < 0.001; continuous NRI = 0.559, P < 0.001; IDI = 0.111, 95%CI: 0.068-0.154, P < 0.001). Similar improvements were observed vs M-APCS in the training cohort (categorical NRI = 0.171, P < 0.001; continuous NRI = 0.460, P < 0.001; IDI = 0.107, P < 0.001) and the validation cohort (categorical NRI = 0.175, P = 0.016; continuous NRI = 0.623, P < 0.001; IDI = 0.117, P < 0.001).
To evaluate model performance in a clinically relevant high-risk population, we extracted a subgroup with APCS score ≥ 4 from the overall cohort (n = 213). This high-risk dataset was derived from the predefined training and validation cohorts, yielding a high-risk training cohort of 149 participants (NAA, n = 68; AA, n = 81) and a high-risk validation cohort of 64 participants (NAA, n = 28; AA, n = 36) (Figure 7).
Baseline characteristics of participants with and without AA in the APCS score ≥ 4 subgroup are summarized in Supplementary Table 3. Overall, demographic and exposure profiles were broadly comparable between AA and NAA participants within each high-risk cohort, and the distribution of APCS and M-APCS scores was similar, consistent with the restricted score range imposed by the APCS score ≥ 4 definition (Supplementary Table 3).
In the APCS score ≥ 4 high-risk subgroup, the tongue image-clinical nomogram showed moderate discrimination in the training set (AUC = 0.752) and comparable discriminative performance in the validation set (AUC = 0.762) (Supplemen
In the APCS score ≥ 4 high-risk subgroup, the tongue image-clinical nomogram and the tongue-only model showed better discrimination than conventional risk scores (Figure 8A and B). In the high-risk training cohort, the nomogram achieved an AUC of 0.752 (95%CI: 0.672-0.831), followed by the tongue-only model (0.734, 95%CI: 0.652-0.816), whereas discrimination for APCS score (0.528, 95%CI: 0.459-0.596), and M-APCS score (0.569, 95%CI: 0.483-0.656) was close to chance (Figure 8A). In the high-risk validation cohort, the tongue-only model showed the highest AUC (0.769, 95%CI: 0.655-0.884), with the nomogram showing comparable performance (0.762, 95%CI: 0.644-0.881); APCS score (0.515, 95%CI: 0.404-0.627), and M-APCS score (0.569, 95%CI: 0.432-0.706) again demonstrated limited discrimination (Figure 8B). DCA suggested that the nomogram and tongue-only model provided greater net benefit than APCS and M-APCS scores across a clinically relevant range of threshold probabilities in both high-risk cohorts (Figure 8C and D). Calibration plots indicated overall acceptable agreement for the tongue-based models, the APCS calibration curve was not shown because the APCS score ≥ 4 restriction yielded a truncated, discrete score range, and M-APCS score showed less stable calibration in this high-risk setting (Figure 8E and F).
To further quantify incremental classification performance within the APCS score ≥ 4 subgroup, we calculated NRI and IDI for the tongue image-clinical nomogram against the conventional scores. Compared with the APCS score, the no
A single cutoff for each method was determined in the APCS score ≥ 4 training subset to target approximately 80% specificity and then applied unchanged to the validation subset (Table 2). The tongue image-clinical nomogram cutoff was 0.5876, achieving a training-set specificity of 79.41% and sensitivity of 66.67%. When applied to the validation cohort, the nomogram yielded sensitivity/specificity of 58.33%/75.0%, with PPV 75.0%, NPV 58.33%, and accuracy 65.62%. For the tongue-only model, the training-derived cutoff (0.6344) produced a training-set specificity of 77.94% and sensitivity of 61.73%, and in validation achieved sensitivity/specificity of 63.89%/82.14% (PPV 82.14%; NPV 63.89%; accuracy 71.88%). In contrast, under the same fixed-specificity strategy, APCS (cutoff = 5) and M-APCS (cutoff = 5) scores showed substantially lower sensitivities in the validation cohort (19.44% and 25.00%, respectively), despite maintaining relatively high specificities (82.14% and 92.86%, respectively). Overall, at comparable false-positive rates, the tongue-based models provided markedly improved case detection for AA in the APCS score ≥ 4 high-risk population (Table 2).
| Method | Cohort | Threshold | TP | FP | TN | FN | Sensitivity (%) | Specificity (%) | PPV (%) | NPV (%) | Accuracy (%) |
| Tongue image-clinical nomogram | Training | 0.5876 | 54 | 14 | 54 | 27 | 66.67 | 79.41 | 79.41 | 66.67 | 72.48 |
| Validation | 0.5876 | 21 | 7 | 21 | 15 | 58.33 | 75.00 | 75.00 | 58.33 | 65.62 | |
| Tongue-only model | Training | 0.6344 | 50 | 15 | 53 | 31 | 61.73 | 77.94 | 76.92 | 63.10 | 69.13 |
| Validation | 0.6344 | 23 | 5 | 23 | 13 | 63.89 | 82.14 | 82.14 | 63.89 | 71.88 | |
| APCS score | Training | 5 | 24 | 16 | 52 | 57 | 29.63 | 76.47 | 60.00 | 47.71 | 51.01 |
| Validation | 5 | 7 | 5 | 23 | 29 | 19.44 | 82.14 | 58.33 | 44.23 | 46.88 | |
| M-APCS score | Training | 5 | 21 | 14 | 54 | 60 | 25.93 | 79.41 | 60.00 | 47.37 | 50.34 |
| Validation | 5 | 9 | 2 | 26 | 27 | 25.00 | 92.86 | 81.82 | 49.06 | 54.69 |
In this prospective observational study of 880 adults undergoing colonoscopy, with standardized tongue imaging and blinded outcome assessment, we developed and internally validated a tongue image-clinical nomogram for AA detection. By combining computerized tongue image analysis with routinely collected clinical variables, the model converted a traditionally subjective visual assessment into a structured and quantifiable approach to AA risk estimation. In the overall cohort, the nomogram showed moderate discrimination (AUC = 0.739 in the training cohort and 0.711 in the validation cohort) and consistently ranked above the tongue-only model, APCS score, and M-APCS score. A more notable finding emerged in the prespecified APCS score ≥ 4 subgroup, where the tongue-based models still separated AA from NAA reasonably well, while the conventional scores showed little discriminatory value. Under a training-derived threshold targeting approximately 80% specificity, the nomogram detected 58.33% of AA cases in the high-risk validation subset, compared with 19.44% for the APCS score and 25.00% for the M-APCS score. These findings suggest that quantitative tongue-image features may provide additional risk information beyond conventional demographic-based scores.
The main clinical value of the model is less its modest improvement in AUC in the overall cohort and more its possible use as an added triage step in opportunistic screening. Questionnaire-based scores such as the APCS and M-APCS can identify individuals at increased baseline risk[7,8], yet they provide limited resolution within those already labeled as high risk. In our study, this ceiling effect was evident after APCS-based enrichment. APCS and M-APCS discrimination declined markedly, while the tongue image-clinical model maintained AUCs of 0.752-0.762. This finding may matter in practice because colonoscopy resources are limited, and clinicians often need better ways to identify which patients within an already high-risk group should be prioritized for colonoscopic evaluation. The difference between the nomo
Given the 157 candidate tongue-image variables, model development emphasized parsimony and interpretability. The modest decline in discrimination from the training to the validation cohort, together with similar Brier scores, acceptable bootstrap-corrected calibration, significant NRI/IDI values, and favorable decision-curve profiles, supports reasonable internal stability and suggests that the model improves individual risk ranking compared with conventional scores. However, these internal metrics do not preclude optimism in a single-center model derived from a high-dimensional feature space, although external validation remains essential before broader clinical adoption. The model in this study should be viewed as an internally validated clinical prediction model that requires external validation before broader clinical use.
Clinically, the tongue image-clinical nomogram should be viewed as a complementary noninvasive adjunct rather than a replacement for established stool-based screening tools. FIT and multitarget stool DNA assays occupy established roles in CRC screening[4,6,9], whereas our model offers a different type of information: Rapid, contactless oral phenotyping combined with routine clinical factors. This use may be most relevant in outpatient visits, health examinations, or pre-colonoscopy triage, where a low-burden adjunct could further refine risk among APCS-defined higher-risk individuals. Because no biospecimen handling is required, oral imaging may also be attractive for patients reluctant to complete stool-based tests, although this potential acceptability advantage was not formally assessed in our cohort. Because FIT results and stool DNA were not collected in this study, no direct head-to-head comparison can be made, and the nomogram should not be presented as superior to or substitutive for these established tools. Although image acquisition is brief, and the prediction process can be automated, claims regarding cost-effectiveness, large-scale dissemination, or smartphone deployment would be premature before cross-device validation and real-world implementation studies.
One distinctive aspect of this study lies in the way the problem was approached methodologically. Computerized tongue image analysis can transform qualitative tongue inspection into standardized quantitative image descriptors[21,22]. Prior tongue-image studies in oncology have largely focused on established cancers, especially gastric cancer[23,24], whereas our study extends quantitative oral phenotyping to a premalignant colorectal endpoint with colonoscopy-confirmed ground truth. The study does not require gastroenterologists to adopt TCM theory as a clinical decision system. Rather, it uses standardized imaging, automated feature extraction, and colonoscopy-confirmed outcomes to convert a historically subjective visual examination into quantitative digital biomarkers. The contribution of the study is not merely applying artificial intelligence to TCM, but developing a multimodal, interpretable, noninvasive tool for a clinically important premalignant colorectal endpoint. This keeps the focus on objective risk assessment rather than theoretical interpretation.
The biological interpretation of the selected tongue-image features should remain cautious. The retained predictors may reflect aspects of oral biofilm ecology[14,26]. Existing work on the oral-gut axis and oral microbial signatures in colorectal neoplasia provides a biologically reasonable context for these associations[12,13,15-17]. Studies linking oral dysbiosis and Fusobacterium nucleatum to colorectal carcinogenesis further support this rationale[27,28]. Similarly, central tongue-body brightness may reflect mucosal light absorption and microcirculatory variation[18,29,30], while low-grade systemic inflammation and early vascular remodeling may provide a broader physiological context for these associations[19,20]. These features should be interpreted as associative digital biomarkers rather than direct mechanistic surrogates of adenoma progression.
Several features of the study strengthen confidence in the observed signal, including prospective recruitment, colonoscopy with histopathological confirmation as the reference standard, blinding of endoscopists, standardized image acquisition under controlled illumination, and automated feature extraction. Several limitations also deserve emphasis. First, this was a single-center colonoscopy-based cohort from a TCM-affiliated hospital and should be regarded as an opportunistic/referral population rather than a community screening sample; spectrum bias and referral bias may therefore influence transportability. Second, regional diet, oral-health behaviors, microbiome composition, and related cultural factors may affect tongue phenotype. Third, although internal consistency was strengthened by the use of one calibrated device and one trained operator, reproducibility across other devices and operators remains unknown. Fourth, the APCS score ≥ 4 validation subset was relatively small, limiting the precision of threshold-based estimates. Finally, residual confounding from oral hygiene, recent illness, salivary factors, and other short-term exposures cannot be fully excluded despite standardized preimaging instructions. The cross-sectional nature of the present analysis also precludes temporal inference regarding whether these tongue phenotypes precede AA development. Future multicenter external validation, device harmonization studies, and mechanistic work integrating oral microbiome or metabolomic profiling will be important to clarify the role of digital tongue phenotyping in AA risk refinement. Overall, these findings suggest that digital tongue phenotyping may serve as a low-burden adjunct for AA risk stratification in opportunistic settings, particularly among individuals already classified as high risk.
In this single-center cohort with internal validation, the noninvasive tongue image–clinical nomogram showed higher discrimination for AA than the APCS and M-APCS scores. Importantly, within the APCS score ≥ 4 subgroup, the nomogram showed higher sensitivity at matched specificity, indicating that it may complement existing risk scores for opportunistic screening and risk refinement in high-risk individuals.
| 1. | Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, Jemal A. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74:229-263. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 16785] [Cited by in RCA: 16746] [Article Influence: 8373.0] [Reference Citation Analysis (31)] |
| 2. | Fearon ER, Vogelstein B. A genetic model for colorectal tumorigenesis. Cell. 1990;61:759-767. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 9134] [Cited by in RCA: 7993] [Article Influence: 222.0] [Reference Citation Analysis (24)] |
| 3. | Corley DA, Jensen CD, Marks AR, Zhao WK, Lee JK, Doubeni CA, Zauber AG, de Boer J, Fireman BH, Schottinger JE, Quinn VP, Ghai NR, Levin TR, Quesenberry CP. Adenoma detection rate and risk of colorectal cancer and death. N Engl J Med. 2014;370:1298-1306. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1862] [Cited by in RCA: 1725] [Article Influence: 143.8] [Reference Citation Analysis (12)] |
| 4. | US Preventive Services Task Force, Davidson KW, Barry MJ, Mangione CM, Cabana M, Caughey AB, Davis EM, Donahue KE, Doubeni CA, Krist AH, Kubik M, Li L, Ogedegbe G, Owens DK, Pbert L, Silverstein M, Stevermer J, Tseng CW, Wong JB. Screening for Colorectal Cancer: US Preventive Services Task Force Recommendation Statement. JAMA. 2021;325:1965-1977. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1663] [Cited by in RCA: 1549] [Article Influence: 309.8] [Reference Citation Analysis (5)] |
| 5. | Hull MA, Rees CJ, Sharp L, Koo S. A risk-stratified approach to colorectal cancer prevention and diagnosis. Nat Rev Gastroenterol Hepatol. 2020;17:773-780. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 84] [Cited by in RCA: 109] [Article Influence: 18.2] [Reference Citation Analysis (4)] |
| 6. | Imperiale TF, Gruber RN, Stump TE, Emmett TW, Monahan PO. Performance Characteristics of Fecal Immunochemical Tests for Colorectal Cancer and Advanced Adenomatous Polyps: A Systematic Review and Meta-analysis. Ann Intern Med. 2019;170:319-329. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 101] [Cited by in RCA: 193] [Article Influence: 27.6] [Reference Citation Analysis (3)] |
| 7. | Yeoh KG, Ho KY, Chiu HM, Zhu F, Ching JY, Wu DC, Matsuda T, Byeon JS, Lee SK, Goh KL, Sollano J, Rerknimitr R, Leong R, Tsoi K, Lin JT, Sung JJ; Asia-Pacific Working Group on Colorectal Cancer. The Asia-Pacific Colorectal Screening score: a validated tool that stratifies risk for colorectal advanced neoplasia in asymptomatic Asian subjects. Gut. 2011;60:1236-1241. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 281] [Cited by in RCA: 263] [Article Influence: 17.5] [Reference Citation Analysis (1)] |
| 8. | Sung JJY, Wong MCS, Lam TYT, Tsoi KKF, Chan VCW, Cheung W, Ching JYL. A modified colorectal screening score for prediction of advanced neoplasia: A prospective study of 5744 subjects. J Gastroenterol Hepatol. 2018;33:187-194. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 68] [Cited by in RCA: 66] [Article Influence: 8.3] [Reference Citation Analysis (0)] |
| 9. | Imperiale TF, Ransohoff DF, Itzkowitz SH, Levin TR, Lavin P, Lidgard GP, Ahlquist DA, Berger BM. Multitarget stool DNA testing for colorectal-cancer screening. N Engl J Med. 2014;370:1287-1297. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1503] [Cited by in RCA: 1357] [Article Influence: 113.1] [Reference Citation Analysis (3)] |
| 10. | Jiang B, Liang X, Chen Y, Ma T, Liu L, Li J, Jiang R, Chen T, Zhang X, Li S. Integrating next-generation sequencing and traditional tongue diagnosis to determine tongue coating microbiome. Sci Rep. 2012;2:936. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 81] [Cited by in RCA: 104] [Article Influence: 7.4] [Reference Citation Analysis (0)] |
| 11. | Maciocia G. Tongue diagnosis in Chinese medicine. 3rd ed. Seattle: Eastland Press, 1995: 119-138. |
| 12. | Kitamoto S, Nagao-Kitamoto H, Hein R, Schmidt TM, Kamada N. The Bacterial Connection between the Oral Cavity and the Gut Diseases. J Dent Res. 2020;99:1021-1029. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 376] [Cited by in RCA: 317] [Article Influence: 52.8] [Reference Citation Analysis (0)] |
| 13. | Flemer B, Warren RD, Barrett MP, Cisek K, Das A, Jeffery IB, Hurley E, O'Riordain M, Shanahan F, O'Toole PW. The oral microbiota in colorectal cancer is distinctive and predictive. Gut. 2018;67:1454-1463. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 534] [Cited by in RCA: 482] [Article Influence: 60.3] [Reference Citation Analysis (4)] |
| 14. | Seerangaiyan K, Jüch F, Winkel EG. Tongue coating: its characteristics and role in intra-oral halitosis and general health-a review. J Breath Res. 2018;12:034001. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 33] [Cited by in RCA: 59] [Article Influence: 7.4] [Reference Citation Analysis (0)] |
| 15. | Zepeda-Rivera M, Minot SS, Bouzek H, Wu H, Blanco-Míguez A, Manghi P, Jones DS, LaCourse KD, Wu Y, McMahon EF, Park SN, Lim YK, Kempchinsky AG, Willis AD, Cotton SL, Yost SC, Sicinska E, Kook JK, Dewhirst FE, Segata N, Bullman S, Johnston CD. A distinct Fusobacterium nucleatum clade dominates the colorectal cancer niche. Nature. 2024;628:424-432. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 339] [Cited by in RCA: 333] [Article Influence: 166.5] [Reference Citation Analysis (0)] |
| 16. | Komiya Y, Shimomura Y, Higurashi T, Sugi Y, Arimoto J, Umezawa S, Uchiyama S, Matsumoto M, Nakajima A. Patients with colorectal cancer have identical strains of Fusobacterium nucleatum in their colorectal cancer and oral cavity. Gut. 2019;68:1335-1337. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 124] [Cited by in RCA: 239] [Article Influence: 34.1] [Reference Citation Analysis (5)] |
| 17. | Bullman S, Pedamallu CS, Sicinska E, Clancy TE, Zhang X, Cai D, Neuberg D, Huang K, Guevara F, Nelson T, Chipashvili O, Hagan T, Walker M, Ramachandran A, Diosdado B, Serna G, Mulet N, Landolfi S, Ramon Y Cajal S, Fasani R, Aguirre AJ, Ng K, Élez E, Ogino S, Tabernero J, Fuchs CS, Hahn WC, Nuciforo P, Meyerson M. Analysis of Fusobacterium persistence and antibiotic response in colorectal cancer. Science. 2017;358:1443-1448. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1345] [Cited by in RCA: 1239] [Article Influence: 137.7] [Reference Citation Analysis (6)] |
| 18. | Ince C, Boerma EC, Cecconi M, De Backer D, Shapiro NI, Duranteau J, Pinsky MR, Artigas A, Teboul JL, Reiss IKM, Aldecoa C, Hutchings SD, Donati A, Maggiorini M, Taccone FS, Hernandez G, Payen D, Tibboel D, Martin DS, Zarbock A, Monnet X, Dubin A, Bakker J, Vincent JL, Scheeren TWL; Cardiovascular Dynamics Section of the ESICM. Second consensus on the assessment of sublingual microcirculation in critically ill patients: results from a task force of the European Society of Intensive Care Medicine. Intensive Care Med. 2018;44:281-299. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 373] [Cited by in RCA: 349] [Article Influence: 43.6] [Reference Citation Analysis (0)] |
| 19. | Zhang X, Liu S, Zhou Y. Circulating levels of C-reactive protein, interleukin-6 and tumor necrosis factor-α and risk of colorectal adenomas: a meta-analysis. Oncotarget. 2016;7:64371-64379. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 13] [Cited by in RCA: 13] [Article Influence: 1.3] [Reference Citation Analysis (0)] |
| 20. | Staton CA, Chetwood AS, Cameron IC, Cross SS, Brown NJ, Reed MW. The angiogenic switch occurs at the adenoma stage of the adenoma carcinoma sequence in colorectal cancer. Gut. 2007;56:1426-1432. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 70] [Cited by in RCA: 69] [Article Influence: 3.6] [Reference Citation Analysis (0)] |
| 21. | Xie J, Jing C, Zhang Z, Xu J, Duan Y, Xu D. Digital tongue image analyses for health assessment. Med Rev (2021). 2021;1:172-198. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 7] [Cited by in RCA: 21] [Article Influence: 4.2] [Reference Citation Analysis (0)] |
| 22. | Lin H, Ning Z, Zhang C, Men S, Zhang D. Computerized tongue image analysis for non-invasive disease screening: a review. Chin Med. 2025;20:196. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 5] [Reference Citation Analysis (0)] |
| 23. | Yuan L, Yang L, Zhang S, Xu Z, Qin J, Shi Y, Yu P, Wang Y, Bao Z, Xia Y, Sun J, He W, Chen T, Chen X, Hu C, Zhang Y, Dong C, Zhao P, Wang Y, Jiang N, Lv B, Xue Y, Jiao B, Gao H, Chai K, Li J, Wang H, Wang X, Guan X, Liu X, Zhao G, Zheng Z, Yan J, Yu H, Chen L, Ye Z, You H, Bao Y, Cheng X, Zhao P, Wang L, Zeng W, Tian Y, Chen M, You Y, Yuan G, Ruan H, Gao X, Xu J, Xu H, Du L, Zhang S, Fu H, Cheng X. Development of a tongue image-based machine learning tool for the diagnosis of gastric cancer: a prospective multicentre clinical cohort study. EClinicalMedicine. 2023;57:101834. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 78] [Cited by in RCA: 77] [Article Influence: 25.7] [Reference Citation Analysis (0)] |
| 24. | Zhu X, Ma Y, Guo D, Men J, Xue C, Cao X, Zhang Z. A Framework to Predict Gastric Cancer Based on Tongue Features and Deep Learning. Micromachines (Basel). 2022;14:53. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1] [Cited by in RCA: 14] [Article Influence: 3.5] [Reference Citation Analysis (0)] |
| 25. | Gupta S, Lieberman D, Anderson JC, Burke CA, Dominitz JA, Kaltenbach T, Robertson DJ, Shaukat A, Syngal S, Rex DK. Recommendations for Follow-Up After Colonoscopy and Polypectomy: A Consensus Update by the US Multi-Society Task Force on Colorectal Cancer. Gastrointest Endosc. 2020;91:463-485.e5. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 265] [Cited by in RCA: 258] [Article Influence: 43.0] [Reference Citation Analysis (7)] |
| 26. | Roldán S, Herrera D, Sanz M. Biofilms and the tongue: therapeutical approaches for the control of halitosis. Clin Oral Investig. 2003;7:189-197. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 71] [Cited by in RCA: 59] [Article Influence: 2.6] [Reference Citation Analysis (0)] |
| 27. | Zhang S, Kong C, Yang Y, Cai S, Li X, Cai G, Ma Y. Human oral microbiome dysbiosis as a novel non-invasive biomarker in detection of colorectal cancer. Theranostics. 2020;10:11595-11606. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 98] [Cited by in RCA: 83] [Article Influence: 13.8] [Reference Citation Analysis (1)] |
| 28. | Rubinstein MR, Wang X, Liu W, Hao Y, Cai G, Han YW. Fusobacterium nucleatum promotes colorectal carcinogenesis by modulating E-cadherin/β-catenin signaling via its FadA adhesin. Cell Host Microbe. 2013;14:195-206. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 2107] [Cited by in RCA: 1936] [Article Influence: 148.9] [Reference Citation Analysis (30)] |
| 29. | Kouadio AA, Jordana F, Koffi NJ, Le Bars P, Soueidan A. The use of laser Doppler flowmetry to evaluate oral soft tissue blood flow in humans: A review. Arch Oral Biol. 2018;86:58-71. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 20] [Cited by in RCA: 28] [Article Influence: 3.1] [Reference Citation Analysis (0)] |
| 30. | Liu M, Zhao J, Lu X, Li G, Wu T, Zhang L. Blood hyperviscosity identification with reflective spectroscopy of tongue tip based on principal component analysis combining artificial neural network. Biomed Eng Online. 2018;17:60. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1] [Cited by in RCA: 2] [Article Influence: 0.3] [Reference Citation Analysis (0)] |