Published online Jul 15, 2026. doi: 10.4251/wjgo.v18.i7.118867
Revised: February 26, 2026
Accepted: March 31, 2026
Published online: July 15, 2026
Processing time: 181 Days and 6.9 Hours
Accurate preoperative identification of advanced (≥ pT3) gastric cancer is essential for selecting candidates for neoadjuvant therapy and planning operative strategy. Endoscopic ultrasonography (EUS) provides routine categorical staging (uT) and quantitative lesion thickness, yet thickness is often simplified using categorical cut-offs, potentially obscuring non-linear risk patterns. We hypothesized that modelling EUS-measured thickness as a continuous predictor could improve individualised estimation of ≥ pT3 risk and provide a probabilistic complement to uT staging.
To develop and validate a preoperative model using EUS lesion thickness to predict pathological deep invasion in gastric cancer.
We retrospectively studied 518 gastrectomy patients (2017-2020) without neo
Deep invasion occurred in 354/518 patients. Discrimination, measured by the area under the receiver operating characteristic curve (AUC), was 0.851 [95% confidence interval (CI): 0.817-0.885]. The model outperformed a uT-only model (AUC 0.810, 95%CI: 0.768-0.851) and a uT ≥ 3 rule (AUC 0.761, 95%CI: 0.721-0.802); the uT ≥ 3 rule had poorer accuracy (Brier score 0.214). In temporal validation, AUC was 0.845 (95%CI: 0.798-0.893) with calibration drift (intercept -0.466; slope 0.623).
An EUS thickness-based model estimates deep invasion risk and may complement routine uT staging. Temporal testing showed preserved discrimination but calibration drift, supporting future recalibration and external va
Core Tip: Preoperative identification of deep gastric wall invasion remains challenging. Endoscopic ultrasonography provides both a categorical T stage and a measurable lesion thickness. We developed and internally validated a logistic model that treats lesion thickness as a continuous predictor, together with routinely available clinical variables, to estimate individualised risk of pathological ≥ pT3 disease. The model showed good discrimination and decision-curve utility, and outperformed uT-based benchmarks. Temporal testing in a later same-institution cohort preserved discrimination but showed calibration drift, supporting future recalibration, external validation, and context-specific threshold selection.
- Citation: Cao MY, Chen YF, Liao YX, Yang Q, Lin SY, Li Y, Weng J, Huang XX, Gao XY, Wang GB, Shan HB, Li JJ. Endoscopic ultrasonography lesion thickness predicts deep invasion in gastric cancer: Development and temporal validation of a preoperative model. World J Gastrointest Oncol 2026; 18(7): 118867
- URL: https://www.wjgnet.com/1948-5204/full/v18/i7/118867.htm
- DOI: https://dx.doi.org/10.4251/wjgo.v18.i7.118867
Gastric cancer remains a major global health burden and is among the leading causes of cancer-related death worldwide despite advances in screening, surgery, and systemic therapy[1]. For patients with potentially curable disease, the depth of tumour invasion is central to tumor-node-metastasis staging, prognostication, and treatment selection. In particular, the distinction between mucosal or submucosal lesions and deeply invasive tumours (pathological T3/T4, ≥ pT3) has major implications for pre-treatment stratification and multidisciplinary treatment planning, including intensified preoperative staging, consideration of neoadjuvant therapy, and operative strategy[2,3].
Endoscopic ultrasonography (EUS) provides high-resolution, layer-by-layer visualisation of the gastric wall and has long been used for local T staging[4,5]. Meta-analyses and large series suggest that EUS performs reasonably well in distinguishing early from advanced gastric cancer, but its accuracy for detailed T1-T4 staging is only moderate and highly heterogeneous[4-7]. Importantly, categorical uT assignment remains a cornerstone of routine preoperative staging, yet its performance may vary across centres and operators and may be influenced by lesion-related factors such as ulceration, fibrosis, tumour size, and diffuse-type histology[5,8-10]. Over- and under-staging are particularly common in these settings and may contribute to misclassification around treatment-relevant boundaries, thereby affecting pre-treatment risk stratification and planning[8-13]. Computed tomography (CT) remains indispensable for assessing nodal status, distant metastasis, and peritoneal dissemination, but its spatial resolution for mural invasion depth is limited; therefore, a complementary, more standardisable quantitative approach may help support individualised pre-treatment risk esti
Beyond qualitative assessment of wall layers, EUS allows measurement of lesion thickness, reflecting the degree of tumour-related wall thickening. Although it is intuitive that thicker lesions correspond to deeper invasion, several studies examining EUS-derived wall thickness or fold morphology as markers of advanced disease have relied on empirical cut-offs and dichotomization[18-22]. Such approaches simplify interpretation but discard continuous information and assume a linear relationship that may not exist, which can reduce predictive accuracy and obscure clinically important risk gradients[23]. Because thickness is a routinely reportable, quantitative metric, it has the potential to serve as a practical adjunct to uT staging, particularly near key decision thresholds or in settings where categorical uT assessment may be less consistent. In addition, most prior studies did not explicitly distinguish between explanatory analyses linking thickness to pathological T stage and true preoperative prediction models that rely solely on information available before surgery. Only a limited number of investigations have examined EUS-derived wall thickness using more flexible modelling approaches rather than empirical cut-offs[19]. Contemporary prediction research standards emphasise modelling continuous predictors without arbitrary dichotomisation, assessing both discrimination and calibration, and considering clinical utility[23-25].
Parallel developments in prediction research have produced numerous nomograms for lymph node metastasis, occult peritoneal dissemination, and prognosis in gastric cancer using clinicopathological variables, CT-derived features, radiomics, and blood-based markers[26-31]. However, most available tools do not incorporate quantitative EUS measurements and seldom focus on estimating the probability of deep mural invasion (≥ pT3), a treatment-relevant boundary that can prompt intensified staging (e.g., targeted imaging or diagnostic laparoscopy where appropriate), multidisciplinary discussion, and consideration of neoadjuvant therapy and operative planning. In addition, several models rely on postoperative information or report discrimination alone without adequate calibration assessment, internal validation, or evaluation of clinical utility, limiting their interpretability and portability for pre-treatment deci
Therefore, we developed and internally validated an EUS-based pre-treatment prediction model centred on continuously modelled lesion thickness and a small set of routinely available clinical variables to estimate individualised risk of pathological ≥ pT3 gastric cancer. To clarify its role in practice, we benchmarked the model against conventional uT-based approaches (a uT-only model and a pragmatic uT ≥ 3 high-risk rule), positioning the output as a probabilistic adjunct to routine staging rather than a replacement. We further examined temporal performance in a subsequent same-institution cohort to provide an initial assessment of model transport over time.
This retrospective study was conducted at Sun Yat-sen University Cancer Center. Consecutive patients with gastric cancer who underwent preoperative EUS and curative-intent gastrectomy between January 2017 and December 2020 were identified from institutional databases. This study was approved by the Institutional Review Board of Sun Yat-sen University Cancer Center, No. SL-B2025-207-01.
Inclusion criteria: (1) Histologically confirmed gastric adenocarcinoma on endoscopic biopsy; (2) Preoperative EUS performed at our institution; (3) No preoperative anticancer therapy; and (4) Complete postoperative pathological sta
Exclusion criteria: (1) Gastro-oesophageal junction carcinoma; (2) Indeterminate pathological diagnosis; (3) Incomplete EUS reports, particularly missing lesion thickness; and (4) Receipt of neoadjuvant therapy before EUS. Patients receiving neoadjuvant chemotherapy were strictly excluded to ensure the reliability of the pathological reference standard (ground truth). Since chemotherapy induces tumour regression and fibrosis (downstaging), it alters the intrinsic correlation between preoperative morphologic features (thickness) and the native depth of invasion. Therefore, a treatment-naive cohort was strictly required to establish the baseline quantitative relationship between lesion thickness and mural invasion. After applying these criteria, 518 of 600 screened patients were included in the final development cohort (2017-2020).
To perform a temporal validation, we split the study population according to calendar time at the same institution. Patients treated during 2017-2020 were assigned to the development cohort (model derivation), and patients treated during 2021-2025 were assigned to the temporal validation cohort. This non-overlapping time-based split was prespecified to evaluate model transportability across a later clinical period while avoiding random-split optimism. The institutional review board approved the study and waived informed consent owing to the retrospective design and use of anonymised data. The study was designed and reported in accordance with the TRIPOD statement[32]. A detailed flowchart of patient screening, exclusion, and formation of the analytic cohort is shown in Figure 1.
All patients underwent diagnostic upper endoscopy followed by EUS after at least 12 hours of fasting. Conventional endoscopy was used to record tumour location, size, and macroscopic appearance. EUS examinations were performed by experienced endosonographers using the same linear-array echoendoscope platform throughout the study period (Olympus GF-UCT260, 5-12 MHz), with a water-filled balloon used when needed for acoustic coupling. Lesion thickness was defined as the maximal vertical distance from the mucosal surface to the serosal outer border at the tumour site on the largest cross-sectional image and was recorded in millimetres. Thickness was measured in real time during the index EUS examination and documented in the formal EUS report; no retrospective re-measurement from stored images was performed. To reduce measurement variability, a standardized departmental protocol was used (minimizing probe compression, maintaining adequate gastric distension, and applying a consistent measurement/reporting workflow). EUS-based uT/uN categories and tumour location were also recorded, and Lauren classification was obtained from preoperative biopsy histology.
All patients underwent gastrectomy with lymphadenectomy according to institutional protocols and contemporary guidelines. Resected specimens were examined by dedicated gastrointestinal pathologists. Pathological T and N categories and overall stage were assigned according to the 8th edition of the American Joint Committee on Cancer/Union for International Cancer Control tumor-node-metastasis classification[25]. Histological subtype and grade, Lauren classification, Borrmann type, lymphovascular invasion, perineural invasion, and margin status were documented. These postoperative variables were summarised descriptively and used in secondary analyses but were not included as predictors in the main preoperative model.
For each patient, we collected age, sex, EUS and endoscopic features (tumour location, Borrmann type, uT and uN categories, maximum EUS-measured lesion thickness), biopsy histology (differentiation grade, Lauren classification), and preoperative laboratory data [serum carcinoembryonic antigen (CEA), carbohydrate antigen 19-9 (CA19-9), carbohydrate antigen 72-4, and human epidermal growth factor receptor-2 and Epstein-Barr virus-encoded RNA status when avai
The primary outcome was pathological T stage dichotomised as ≥ pT3 vs T1/T2 on the resection specimen, reflecting a clinically relevant distinction that can inform multidisciplinary discussion, consideration of neoadjuvant therapy, and operative strategy. For the primary prediction model, candidate predictors were restricted to preoperative variables: EUS-measured lesion thickness (mm), age, sex, tumour location, Lauren type, and serum CEA (ng/mL) and CA19-9 (U/mL); carbohydrate antigen 72-4, human epidermal growth factor receptor-2, and Epstein-Barr virus-encoded RNA were used only in descriptive and exploratory analyses.
All analyses were performed using R (version 4.5.2; R Foundation for Statistical Computing, Vienna, Austria). Two-sided P values < 0.05 were considered statistically significant. The primary model was a multivariable logistic regression for pathological ≥ pT3 (vs T1-2). Lesion thickness was entered as a restricted cubic spline with four knots to allow for a non-linear association with the log-odds of ≥ pT3[33]. Other predictors were entered as linear or categorical terms. Wald tests were used to assess the overall and non-linear effects of thickness and other covariates.
Model performance was summarised by the C-statistic [area under the receiver operating characteristic curve (AUC)] with 95% confidence intervals (CIs), the Brier score, and Nagelkerke’s R2. Calibration was evaluated using calibration plots, calibration intercept and slope, and decile-based observed-versus-predicted summaries.
Internal validation used bootstrap resampling (B = 1000) to obtain optimism-corrected estimates of discrimination, calibration, and overall performance[34]. To explore the impact of functional form and additional covariates, we also examined: (1) A thickness-only model using the same spline specification; and (2) A full model in which thickness entered linearly. AUC and Brier scores were compared across models. For benchmarking against routine staging, we also evaluated a uT-only logistic model and a dichotomous rule (uT ≥ 3 as high risk), and compared AUC and Brier scores across these models using the same analytic cohort.
Clinical utility was explored using decision-curve analysis, plotting net benefit for the full and thickness-only models vs “treat-all” and “treat-none” strategies over threshold probabilities from 0.10 to 0.70. In addition, a cut-off based on Youden’s index and thresholds with predefined high sensitivity (approximately 0.90) and high specificity (approximately 0.90) were used to summarise false-negative and false-positive classifications in the cohort[35].
Sensitivity analyses included Firth penalised logistic regression, a mixed-effects model with a random intercept for endoscopists, and a linear-thickness specification. To reduce potential overfitting, we applied uniform shrinkage to the final model using the bootstrap-corrected calibration slope as the shrinkage factor, and then re-estimated the intercept so that the mean predicted risk in the derivation cohort equalled the observed prevalence. Thickness was modelled using restricted cubic splines with four knots at the 5th/35th/65th/95th percentiles (3.691 mm, 8.360 mm, 12.300 mm, and 21.415 mm). For validation and implementation, the model should be applied using the final uniformly shrunk coefficients (Supplementary Table 1), while the unshrunk estimates are reported in Table 1 and the Supplementary material for inference and interpretability.
| Variable1 | OR | 95%CI | P value |
| EUS thickness (overall effect)2 | N/A | N/A | < 0.001 |
| Lauren: Diffuse-type vs intestinal-type | 1.700 | 0.932-3.101 | 0.084 |
| Lauren: Mixed-type vs intestinal-type | 2.315 | 1.257-4.261 | 0.007 |
| Age (per 1-year increase) | 1.002 | 0.982-1.024 | 0.824 |
| Gender: Male vs female | 0.770 | 0.474-1.250 | 0.290 |
| Tumour site: Middle vs upper | 1.066 | 0.554-2.051 | 0.848 |
| Tumour site: Lower vs upper | 1.127 | 0.567-2.240 | 0.733 |
| CEA (per unit increase) | 1.126 | 1.026-1.235 | 0.013 |
| CA19-9 (per unit increase) | 1.009 | 0.997-1.022 | 0.135 |
A complete-case approach was used. Records with missing values in candidate predictors or the outcome were excluded during dataset curation before model fitting. Variable-level missingness summaries for the final analysis datasets are provided in Supplementary Table 2. The final multivariable model was fitted in 518 complete cases in the development cohort, and temporal validation was performed in 246 complete cases. Temporal validation applied the previously developed uniformly shrunk model (Supplementary Table 1; derivation knots retained) to the 2021-2025 cohort without refitting. Performance was assessed by discrimination (AUC with 95%CI), overall accuracy (Brier score), and calibration (calibration intercept and slope).
A total of 764 eligible patients were included, with 518 patients in the development cohort (2017-2020) and 246 patients in the temporal validation cohort (2021-2025). Baseline characteristics of the two cohorts are summarised in Table 2. Compared with the development cohort, the temporal validation cohort showed a lower proportion of ≥ pT3 disease and a lower lesion-thickness distribution overall. The median (interquartile range) lesion thickness was 10.10 mm (7.40-13.60) in the development cohort and 8.00 mm (4.31-12.98) in the temporal validation cohort, and the proportion of ≥ pT3 disease was 68.3% vs 47.2%, respectively.
| Variable1 | Development cohort (2017-2020), n = 518 | Temporal validation cohort (2021-2025), n = 246 | Overall, n = 764 | SMD |
| Age, years | 57.50 (49.00-65.00) | 60.00 (53.00-67.00) | 59.00 (50.00-66.00) | 0.219 |
| Sex | 0.163 | |||
| Male | 285 (55.0) | 155 (63.0) | 440 (57.6) | |
| Female | 233 (45.0) | 91 (37.0) | 324 (42.4) | |
| Tumour location | 0.164 | |||
| Upper stomach (cardia/fundus) | 95 (18.3) | 37 (15.0) | 132 (17.3) | |
| Middle stomach (body/angle) | 242 (46.7) | 108 (43.9) | 350 (45.8) | |
| Lower stomach (antrum/pylorus/remnant) | 181 (34.9) | 101 (41.1) | 282 (36.9) | |
| Lauren classification | 0.191 | |||
| Intestinal | 168 (32.4) | 97 (39.4) | 265 (34.7) | |
| Diffuse | 198 (38.2) | 80 (32.5) | 278 (36.4) | |
| Mixed | 152 (29.3) | 69 (28.0) | 221 (28.9) | |
| EUS lesion thickness, mm | 10.10 (7.40-13.60) | 8.00 (4.31-12.98) | 9.70 (6.17-13.50) | 0.346 |
| CEA | 2.13 (1.37-4.00) | 1.96 (1.19-3.25) | 2.07 (1.28-3.63) | 0.062 |
| CA19-9 | 10.79 (6.15-19.63) | 9.79 (5.45-17.62) | 10.48 (5.98-18.94) | 0.070 |
| Pathological stage | 0.439 | |||
| pT1-2 | 164 (31.7) | 130 (52.8) | 294 (38.5) | |
| ≥ pT3 | 354 (68.3) | 116 (47.2) | 470 (61.5) |
In the multivariable logistic regression model, EUS lesion thickness remained the strongest predictor of pathological ≥ pT3. Thickness showed a highly significant overall association with the outcome (Wald χ2, P < 0.001) and a significant non-linear component (P for non-linearity = 0.005). The non-linear thickness-risk association is illustrated in Figure 2, with a steep increase in predicted risk across the mid-range and attenuation at higher thickness values. Compared with intestinal-type Lauren classification, mixed-type tumours were independently associated with a higher risk of ≥ pT3 [odds ratio (OR) = 2.32, 95%CI: 1.26-4.26; P = 0.007], whereas diffuse-type showed a non-significant trend towards increased risk (OR = 1.70, 95%CI: 0.93-3.10; P = 0.084). Serum CEA was modestly but significantly associated with ≥ pT3 (OR = 1.13 per unit increase, 95%CI: 1.03-1.24; P = 0.013), while CA19-9, age, sex and tumour site were not significantly associated with ≥ pT3 in the multivariable model (Table 1). The complete set of regression coefficients is provided in Supplementary Table 1 (final uniformly shrunk coefficients with re-estimated intercept) to facilitate external validation; the original (unshrunk) regression estimates are provided in Supplementary Table 3 for inference and interpretability.
The apparent AUC of the full spline-based model was 0.851 (95%CI: 0.817-0.885), and the corresponding receiver-operating characteristic curve is presented in Figure 3. The Brier score was 0.142, and Nagelkerke’s R2 was 0.444, indicating good overall performance[36]. Calibration was close to ideal: The apparent calibration intercept and slope were near 0 and 1, respectively. Bootstrap internal validation suggested modest optimism, with an optimism-corrected calibration slope of 0.916. We therefore applied uniform shrinkage using this slope as the shrinkage factor and re-estimated the intercept (final shrunk intercept -4.087), and the final shrunk coefficients are reported in Supplementary Table 1. In the bootstrap validation, the optimism-corrected AUC was 0.838 (from Dxy = 0.675) and the optimism-corrected Brier score was 0.150.
For a representative patient profile, the predicted probability of ≥ pT3 was approximately 0.20 at 5 mm, approximately 0.60 at 10 mm, and > 0.80 at 15 mm, before plateauing above 20 mm. Thus, relatively small changes in thickness within the 5-15 mm range were associated with substantial shifts in estimated risk, whereas further thickening beyond 20-25 mm had a more limited impact on the predicted probability of ≥ pT3 (Figure 2).
The final full spline-based model used for prediction included lesion thickness (modelled using restricted cubic splines), Lauren classification, age, sex, tumour location, CEA, and CA19-9. To improve reproducibility and enable external validation, the complete model specification (including the shrinkage-adjusted regression coefficients, spline basis terms for lesion thickness, and the model intercept) is provided in Supplementary Table 1 and the Supplementary material. The implementation equation is expressed as P = 1/[1 + exp(-LP)], with the linear predictor (LP) defined in the Supplementary material.
In the same-institution temporal validation cohort (2021-2025; n = 246; event rate 47.2%), the model demonstrated good discrimination (AUC = 0.845, 95%CI: 0.798-0.893) with a Brier score of 0.166. Calibration assessment showed modest overestimation of risk (calibration intercept -0.466, 95%CI: -0.795 to -0.136) and some over-dispersion of predictions (calibration slope 0.623, 95%CI: 0.435-0.810). Additional temporal validation metrics are summarized in Supplementary Table 4.
The thickness-only model (restricted cubic spline thickness as the only predictor) had an AUC of 0.825, which was lower than that of the full spline-based model (ΔAUC = 0.026; DeLong P < 0.05). Its Brier score was also higher, indicating poorer overall performance. A full model in which thickness entered linearly produced an AUC of 0.847 (ΔAUC vs spline-based model = 0.004; DeLong P = 0.036). Although this small difference in AUC reached statistical significance, the magnitude of ΔAUC was very small and is unlikely to be clinically meaningful. We therefore retained the spline specification primarily because it offers a more flexible representation of the thickness-risk relationship (Supplementary Table 5).
Decision-curve analysis demonstrated that, over threshold probabilities of approximately 0.10-0.70, both the full model and the thickness-only model provided greater net benefit than default “treat-all” or “treat-none” strategies, and the full model consistently yielded a slightly higher net benefit than the thickness-only model (Figure 3). Net benefit remained positive at higher thresholds (e.g., 0.600), within a region of the risk curve where lesion thickness is clinically informative; in this range, the full model continued to marginally outperform the thickness-only model, although the incremental gain was modest. From a clinical perspective, this pattern suggests that the full model may help support risk-adapted decisions for intensified staging workup and/or consideration of neoadjuvant treatment by improving identification of patients at higher risk of ≥ pT3 disease while limiting unnecessary escalation in lower-risk patients.
Using the full model, a Youden-based “balanced” threshold achieved moderate sensitivity and high specificity, whereas a high-sensitivity threshold (target sensitivity approximately 0.90) substantially reduced the miss rate of ≥ pT3 cancers at the cost of labelling almost half of pT1-T2 cases as high risk. Conversely, a high-specificity threshold (target specificity approximately 0.90) greatly limited overtreatment among pT1-2 patients but allowed a sizeable fraction of ≥ pT3 cases to be missed. These probability thresholds were derived from the receiver-operating characteristic curve of the full model and were used to illustrate different trade-offs between missed ≥ pT3 cases and potential overtreatment. The corresponding classification consequences (true positives, false negatives, false positives, and true negatives) for each strategy are summarised in Supplementary Table 6. Importantly, these thresholds are presented as illustrative decision points for trade-off visualization rather than prescriptive universal clinical cut-offs, and threshold selection should be tailored to local clinical context and resource considerations.
To benchmark the proposed prediction model against routine clinical practice, we compared the full multivariable model with two uT-based benchmarks: (1) A uT-only logistic model; and (2) A simple rule defining high risk as uT ≥ 3. The full model showed higher discrimination than the uT-only model (AUC= 0.851, 95%CI: 0.817-0.885 vs 0.810, 95%CI: 0.768-0.851) and the uT ≥ 3 rule (AUC = 0.761, 95%CI: 0.721-0.802). Overall probability accuracy was similar between the full model and the uT-only model (Brier 0.142 vs 0.142), whereas the uT ≥ 3 rule had substantially worse accuracy (Brier 0.2143), consistent with information loss from dichotomizing stage categories.
Firth penalised logistic regression yielded similar conclusions to the primary model, with slightly lower discrimination (AUC 0.839 vs 0.851), suggesting that small-sample bias or quasi-separation was unlikely to materially change the main inferences. A mixed-effects logistic regression model including a random intercept for individual endoscopists showed negligible between-operator variance, suggesting limited clustering by endoscopist. Finally, the conclusions regarding the importance and shape of the thickness-risk relationship were robust when comparing the spline-based and linear-thickness models.
In this single-centre retrospective cohort, we developed a preoperative prediction model integrating EUS-measured lesion thickness with routinely available clinical variables and found that lesion thickness was the dominant predictor of pathological ≥ pT3 risk. The model showed good discrimination and calibration in the development cohort and main
The key contribution of this study is not merely the use of EUS in preoperative staging, but the transformation of EUS-measured lesion thickness from a categorical impression into a continuous, nonlinearly modelled predictor linked directly to the clinically critical endpoint of ≥ pT3 invasion. Compared with prior prediction models that focused on other outcomes (e.g., lymph node metastasis, peritoneal metastasis, or early-stage depth stratification), our model specifically targets a treatment-relevant decision node and provides individualised risk probabilities rather than a binary stage label.
These findings complement previous work on EUS staging in gastric cancer. Earlier studies and meta-analyses have highlighted the moderate and heterogeneous accuracy of qualitative EUS T staging, particularly in patients with ulcerated lesions, fibrosis, or diffuse-type tumours[4-6,9,10,13]. Several groups have reported that increased gastric wall thickness or prominent folds on EUS are more common in advanced or infiltrative disease[18-22]. However, most prior studies classified thickness using arbitrary cut-offs and primarily focused on explanatory associations with pathological stage rather than preoperative prediction. In contrast, our study models thickness continuously (using restricted cubic splines) and applies it to a clinically oriented preoperative prediction model for ≥ pT3 disease, with evaluation of discrimination, calibration, internal validation, and decision-curve utility[24,32].
From a clinical perspective, the model is positioned as an objective second opinion to refine risk stratification, particularly for patients initially considered for upfront surgery. In our cohort, which consisted of patients clinically judged as candidates for primary resection, 68.3% were eventually confirmed to have pathological ≥ pT3 disease. This high prevalence of advanced pathology highlights a critical limitation of conventional staging: Many patients are effectively under-staged preoperatively. By providing a quantitative probability of deep invasion, our model aims to identify “occult” advanced disease among these patients who might otherwise be scheduled for immediate surgery, thereby reducing the risk of under-treatment (inappropriate upfront surgery) and ensuring that candidates for neoadjuvant therapy are not missed due to subjective under-estimation of T stage. Among patients with resectable disease, the estimated risk of deep mural invasion could also inform discussions about neoadjuvant chemotherapy, particularly when interpreted together with nodal status or other high-risk features. Although thickness around 10-12 mm lay within an informative region of the risk curve in our cohort, clinical decisions should rely on predicted probabilities from the full multivariable model rather than thickness alone. Consistent with this, decision-curve analysis suggested greater net benefit for model-guided stratification than default “treat-all” or “treat-none” strategies across a clinically relevant threshold range, and the thresholds presented here should be interpreted as illustrative rather than prescriptive.
EUS uT category was a strong predictor of pathological advanced invasion in our cohort and remains a clinically important staging descriptor in routine preoperative assessment. In our cohort, the proposed multivariable model provided higher discrimination than a uT-only benchmark and a simple uT ≥ 3 rule, while overall probability accuracy was similar between the full model and uT-only. These findings support positioning the model as an adjunct tool for risk estimation rather than a replacement for standard EUS staging. In practice, the accuracy of uT staging may vary across settings because of operator dependence and differences in equipment and experience. The practical value of this quantitative model should therefore be tested in multicentre cohorts with different platforms, operator experience levels, and case-mix.
Importantly, beyond internal validation, we also evaluated the model in a chronologically later temporal validation cohort (2021-2025) from the same institution. In that setting, discrimination remained acceptable, but calibration drift was observed, with a calibration intercept of -0.466 and a calibration slope of 0.623. Clinically, this indicates that the model still ranked patients reasonably well by risk, but tended to overestimate the absolute probability of ≥ pT3 disease and produced predictions that were too extreme in the later cohort. This pattern is plausible given temporal shifts in case-mix, tumour thickness distribution, and staging or treatment practices, and it underscores that good discrimination alone is insufficient for clinical deployment when absolute risk estimates are used to support treatment decisions. Accordingly, local recalibration (at minimum recalibration-in-the-large, and potentially slope updating) should be considered before routine use in later-period or external populations. At the same time, temporal validation within a single centre cannot substitute for independent external validation across institutions.
In summary, modelling EUS-measured lesion thickness as a continuous non-linear predictor improved individualised estimation of ≥ pT3 risk over simplified approaches. By integrating quantitative EUS thickness with routinely available clinical variables, the model supports pre-treatment risk stratification and highlights the clinical value of quantitative EUS features in gastric cancer care.
Despite these strengths, several limitations should be acknowledged. First, this was a single-centre study conducted in a high-volume specialist setting with a relatively high prevalence of ≥ pT3 disease (68.3%). Model performance may differ in community hospitals, in centres with less experience in EUS, or in regions with a different spectrum of disease. In addition, our cohort included only patients who underwent curative gastrectomy without preoperative therapy. This introduces a selection bias inherent to surgical cohorts, because patients with clearly bulky or unresectable disease would likely proceed directly to systemic therapy. However, this design was intentional. By focusing on the “treatment-naive” population eligible for surgery, we aimed to target the diagnostically difficult grey zone where objective quantitative guidance is most needed, rather than verifying advanced status in obvious cases that pose no diagnostic dilemma. Although the mixed-effects analysis showed minimal between-operator variance, residual operator-dependent diffe
Second, we relied on conventional EUS rather than newer modalities such as contrast-enhanced EUS or elastography, which may further refine pre-treatment risk stratification. The model also did not incorporate CT or magnetic resonance imaging variables, and integrating EUS-derived metrics with cross-sectional imaging might improve predictive perfor
Third, an important limitation is that external validation was not performed. Although we included temporal validation within the same institution (2017-2020 development; 2021-2025 validation), broader generalizability remains uncertain across institutions with different EUS operators, equipment platforms, measurement habits, and patient case-mix. Therefore, large-scale multicenter external validation should be considered the primary next step before routine implementation. In addition, although thickness was measured using a standardized departmental approach, inter-observer and intra-observer variability in EUS-based lesion thickness measurement remains a practical concern, parti
In summary, we developed and internally validated a preoperative prediction model for pathological ≥ pT3 gastric cancer based on EUS-measured lesion thickness and a small set of routinely available clinical variables. The model showed good discrimination (optimism-corrected AUC 0.838), reasonable calibration, and favourable decision-curve profiles, and supports the concept that quantitative EUS thickness carries clinically relevant information about invasion depth. After rigorous external validation and, ideally, prospective impact evaluation, this approach may help refine the selection of patients for neoadjuvant therapy and surgical planning by providing individualised estimates of the probability of advanced pathological T stage, thereby contributing to more tailored management of gastric cancer.
The authors thank the clinicians, endoscopy staff, and pathology staff at Sun Yat-sen University Cancer Center for their support in case identification, data collection, and clinical coordination. We also acknowledge the support of the biostatistics team for methodological discussion and statistical review. We sincerely thank the patients whose anonymised data contributed to this retrospective study.
| 1. | Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71:209-249. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 76817] [Cited by in RCA: 70403] [Article Influence: 14080.6] [Reference Citation Analysis (61)] |
| 2. | Japanese Gastric Cancer Association. Japanese Gastric Cancer Treatment Guidelines 2021 (6th edition). Gastric Cancer. 2023;26:1-25. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 976] [Cited by in RCA: 958] [Article Influence: 319.3] [Reference Citation Analysis (11)] |
| 3. | Lordick F, Carneiro F, Cascinu S, Fleitas T, Haustermans K, Piessen G, Vogel A, Smyth EC; ESMO Guidelines Committee. Gastric cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Ann Oncol. 2022;33:1005-1020. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1159] [Cited by in RCA: 1053] [Article Influence: 263.3] [Reference Citation Analysis (15)] |
| 4. | de Nucci G, Gabbani T, Impellizzeri G, Deiana S, Biancheri P, Ottaviani L, Frazzoni L, Mandelli ED, Soriani P, Vecchi M, Manes G, Manno M. Linear EUS Accuracy in Preoperative Staging of Gastric Cancer: A Retrospective Multicenter Study. Diagnostics (Basel). 2023;13:1842. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 6] [Reference Citation Analysis (0)] |
| 5. | Ziogas DI, Kalakos N, Manolakis A, Voulgaris T, Vezakis I, Tadic M, Papanikolaou IS. Endoscopic Ultrasound (EUS) in Gastric Cancer: Current Applications and Future Perspectives. Diseases. 2025;13:234. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 2] [Reference Citation Analysis (0)] |
| 6. | Mocellin S, Pasquali S. Diagnostic accuracy of endoscopic ultrasonography (EUS) for the preoperative locoregional staging of primary gastric cancer. Cochrane Database Syst Rev. 2015;2015:CD009944. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 91] [Cited by in RCA: 121] [Article Influence: 11.0] [Reference Citation Analysis (2)] |
| 7. | Cardoso R, Coburn N, Seevaratnam R, Sutradhar R, Lourenco LG, Mahar A, Law C, Yong E, Tinmouth J. A systematic review and meta-analysis of the utility of EUS for preoperative staging for gastric cancer. Gastric Cancer. 2012;15 Suppl 1:S19-S26. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 129] [Cited by in RCA: 115] [Article Influence: 8.2] [Reference Citation Analysis (4)] |
| 8. | Lee HH, Lim CH, Park JM, Cho YK, Song KY, Jeon HM, Park CH. Low accuracy of endoscopic ultrasonography for detailed T staging in gastric cancer. World J Surg Oncol. 2012;10:190. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 33] [Cited by in RCA: 37] [Article Influence: 2.6] [Reference Citation Analysis (4)] |
| 9. | Kim SJ, Lim CH, Lee BI. Accuracy of Endoscopic Ultrasonography for Determining the Depth of Invasion in Early Gastric Cancer. Turk J Gastroenterol. 2022;33:785-792. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 11] [Cited by in RCA: 16] [Article Influence: 4.0] [Reference Citation Analysis (2)] |
| 10. | Akashi K, Yanai H, Nishikawa J, Satake M, Fukagawa Y, Okamoto T, Sakaida I. Ulcerous change decreases the accuracy of endoscopic ultrasonography diagnosis for the invasive depth of early gastric cancer. Int J Gastrointest Cancer. 2006;37:133-138. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 13] [Cited by in RCA: 21] [Article Influence: 1.1] [Reference Citation Analysis (0)] |
| 11. | Park JS, Kim H, Bang B, Kwon K, Shin Y. Accuracy of endoscopic ultrasonography for diagnosing ulcerative early gastric cancers. Medicine (Baltimore). 2016;95:e3955. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 26] [Cited by in RCA: 26] [Article Influence: 2.6] [Reference Citation Analysis (0)] |
| 12. | Pei Q, Wang L, Pan J, Ling T, Lv Y, Zou X. Endoscopic ultrasonography for staging depth of invasion in early gastric cancer: A meta-analysis. J Gastroenterol Hepatol. 2015;30:1566-1573. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 45] [Cited by in RCA: 43] [Article Influence: 3.9] [Reference Citation Analysis (0)] |
| 13. | Shi D, Xi XX. Factors Affecting the Accuracy of Endoscopic Ultrasonography in the Diagnosis of Early Gastric Cancer Invasion Depth: A Meta-analysis. Gastroenterol Res Pract. 2019;2019:8241381. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 23] [Cited by in RCA: 21] [Article Influence: 3.0] [Reference Citation Analysis (0)] |
| 14. | Ungureanu BS, Sacerdotianu VM, Turcu-Stiolica A, Cazacu IM, Saftoiu A. Endoscopic Ultrasound vs. Computed Tomography for Gastric Cancer Staging: A Network Meta-Analysis. Diagnostics (Basel). 2021;11:134. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 29] [Cited by in RCA: 29] [Article Influence: 5.8] [Reference Citation Analysis (5)] |
| 15. | Kwee RM, Kwee TC. Imaging in local staging of gastric cancer: a systematic review. J Clin Oncol. 2007;25:2107-2116. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 253] [Cited by in RCA: 221] [Article Influence: 11.6] [Reference Citation Analysis (3)] |
| 16. | Kim SJ, Kim HH, Kim YH, Hwang SH, Lee HS, Park DJ, Kim SY, Lee KH. Peritoneal metastasis: detection with 16- or 64-detector row CT in patients undergoing surgery for gastric cancer. Radiology. 2009;253:407-415. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 134] [Cited by in RCA: 118] [Article Influence: 6.9] [Reference Citation Analysis (1)] |
| 17. | Chen CY, Hsu JS, Wu DC, Kang WY, Hsieh JS, Jaw TS, Wu MT, Liu GC. Gastric cancer: preoperative local staging with 3D multi-detector row CT--correlation with surgical and histopathologic results. Radiology. 2007;242:472-482. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 201] [Cited by in RCA: 203] [Article Influence: 10.7] [Reference Citation Analysis (1)] |
| 18. | Mitsunaga A, Tagata, Hamano, Teramoto, Shirato, Shirato, Nishino. A new method of endoscopic ultrasonography for determining lesion depth in early gastric cancer. Gastrointest Cancer: Targets Ther. 2011;1-10. [DOI] [Full Text] |
| 19. | Seo JY, Kim DH, Ahn JY, Choi KD, Kim HJ, Na HK, Lee JH, Jung KW, Song HJ, Lee GH, Jung HY. Differential Diagnosis of Thickened Gastric Wall between Hypertrophic Gastritis and Borrmann Type 4 Advanced Gastric Cancer. Gut Liver. 2024;18:961-969. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 4] [Cited by in RCA: 4] [Article Influence: 2.0] [Reference Citation Analysis (0)] |
| 20. | Lim H, Lee GH, Na HK, Ahn JY, Lee JH, Choi KS, Kim DH, Choi KD, Song HJ, Jung HY, Kim JH, Kim D, Park YS. Use of Endoscopic Ultrasound to Evaluate Large Gastric Folds: Features Predictive of Malignancy. Ultrasound Med Biol. 2015;41:2614-2620. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 2] [Cited by in RCA: 5] [Article Influence: 0.5] [Reference Citation Analysis (0)] |
| 21. | Agarwala R, Shah J, Dutta U. Thickened Gastric Folds: Approach. J Dig Endosc. 2018;09:149-154. [RCA] [DOI] [Full Text] [Cited by in Crossref: 3] [Cited by in RCA: 6] [Article Influence: 0.9] [Reference Citation Analysis (0)] |
| 22. | Abe S, Oda I, Shimazu T, Kinjo T, Tada K, Sakamoto T, Kusano C, Gotoda T. Depth-predicting score for differentiated early gastric cancer. Gastric Cancer. 2011;14:35-40. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 97] [Cited by in RCA: 83] [Article Influence: 5.5] [Reference Citation Analysis (0)] |
| 23. | Royston P, Altman DG, Sauerbrei W. Dichotomizing continuous predictors in multiple regression: a bad idea. Stat Med. 2006;25:127-141. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1827] [Cited by in RCA: 1567] [Article Influence: 78.4] [Reference Citation Analysis (0)] |
| 24. | Bouwmeester W, Zuithoff NP, Mallett S, Geerlings MI, Vergouwe Y, Steyerberg EW, Altman DG, Moons KG. Reporting and methods in clinical prediction research: a systematic review. PLoS Med. 2012;9:1-12. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 446] [Cited by in RCA: 446] [Article Influence: 31.9] [Reference Citation Analysis (4)] |
| 25. | Amin MB, Greene FL, Edge SB, Compton CC, Gershenwald JE, Brookland RK, Meyer L, Gress DM, Byrd DR, Winchester DP. The Eighth Edition AJCC Cancer Staging Manual: Continuing to build a bridge from a population-based to a more "personalized" approach to cancer staging. CA Cancer J Clin. 2017;67:93-99. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 4716] [Cited by in RCA: 4851] [Article Influence: 539.0] [Reference Citation Analysis (11)] |
| 26. | Dong D, Fang MJ, Tang L, Shan XH, Gao JB, Giganti F, Wang RP, Chen X, Wang XX, Palumbo D, Fu J, Li WC, Li J, Zhong LZ, De Cobelli F, Ji JF, Liu ZY, Tian J. Deep learning radiomic nomogram can predict the number of lymph node metastasis in locally advanced gastric cancer: an international multicenter study. Ann Oncol. 2020;31:912-920. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 360] [Cited by in RCA: 316] [Article Influence: 52.7] [Reference Citation Analysis (6)] |
| 27. | Dong D, Tang L, Li ZY, Fang MJ, Gao JB, Shan XH, Ying XJ, Sun YS, Fu J, Wang XX, Li LM, Li ZH, Zhang DF, Zhang Y, Li ZM, Shan F, Bu ZD, Tian J, Ji JF. Development and validation of an individualized nomogram to identify occult peritoneal metastasis in patients with advanced gastric cancer. Ann Oncol. 2019;30:431-438. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 365] [Cited by in RCA: 338] [Article Influence: 48.3] [Reference Citation Analysis (0)] |
| 28. | Bao D, Yang Z, Chen S, Li K, Hu Y. Construction of a Nomogram Model for Predicting Peritoneal Dissemination in Gastric Cancer Based on Clinicopathologic Features and Preoperative Serum Tumor Markers. Front Oncol. 2022;12:844786. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 8] [Reference Citation Analysis (0)] |
| 29. | Hong Y, Li X, Liu Z, Fu C, Nie M, Chen C, Feng H, Gan S, Zeng Q. Predicting tumor invasion depth in gastric cancer: developing and validating multivariate models incorporating preoperative IVIM-DWI parameters and MRI morphological characteristics. Eur J Med Res. 2024;29:431. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 6] [Reference Citation Analysis (0)] |
| 30. | Zhang Y, Liu Y, Zhang J, Wu X, Ji X, Fu T, Li Z, Wu Q, Bu Z, Ji J. Construction and external validation of a nomogram that predicts lymph node metastasis in early gastric cancer patients using preoperative parameters. Chin J Cancer Res. 2018;30:623-632. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 12] [Cited by in RCA: 13] [Article Influence: 1.6] [Reference Citation Analysis (0)] |
| 31. | Kim SM, Min BH, Ahn JH, Jung SH, An JY, Choi MG, Sohn TS, Bae JM, Kim S, Lee H, Lee JH, Kim YW, Ryu KW, Kim JJ, Lee JH. Nomogram to predict lymph node metastasis in patients with early gastric cancer: a useful clinical tool to reduce gastrectomy after endoscopic resection. Endoscopy. 2020;52:435-443. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 55] [Cited by in RCA: 52] [Article Influence: 8.7] [Reference Citation Analysis (1)] |
| 32. | Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): the TRIPOD statement. J Clin Epidemiol. 2015;68:134-143. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 206] [Cited by in RCA: 245] [Article Influence: 22.3] [Reference Citation Analysis (0)] |
| 33. | Desquilbet L, Mariotti F. Dose-response analyses using restricted cubic spline functions in public health research. Stat Med. 2010;29:1037-1057. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 261] [Cited by in RCA: 1072] [Article Influence: 67.0] [Reference Citation Analysis (2)] |
| 34. | Steyerberg EW, Harrell FE Jr, Borsboom GJ, Eijkemans MJ, Vergouwe Y, Habbema JD. Internal validation of predictive models: efficiency of some procedures for logistic regression analysis. J Clin Epidemiol. 2001;54:774-781. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 2254] [Cited by in RCA: 1997] [Article Influence: 79.9] [Reference Citation Analysis (3)] |
| 35. | Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26:565-574. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 4224] [Cited by in RCA: 4099] [Article Influence: 205.0] [Reference Citation Analysis (5)] |
| 36. | Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21:128-138. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 4068] [Cited by in RCA: 3708] [Article Influence: 231.8] [Reference Citation Analysis (7)] |