TO THE EDITOR
I read with great interest the article by Qian et al[1] published in the World Journal of Gastroenterology. The authors present a valuable study addressing early recurrence (ER) of hepatocellular carcinoma (HCC) after resection in cirrhotic patients - an area of major clinical importance. Their integration of computed tomography (CT)-based radiomic features with clinical variables is commendable, and this work represents a meaningful step toward quantitative imaging–based prognostication in postoperative HCC. Several methodological considerations, however, may enhance the interpretability, reproducibility, and generalizability of radiomics-driven ER prediction. These considerations are aligned with recent international initiatives to improve radiomics reproducibility and clinical readiness, including the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis Or Diagnosis (TRIPOD) + artificial intelligence (AI) statement and the radiomics quality score (RQS) 2.0 framework.
First, clarification of the outcome definition and modeling framework would strengthen the prognostic interpretation of the combined model. Qian et al[1] examined 214 patients with cirrhosis who underwent curative hepatectomy, among whom 114 experienced ER, and developed a machine-learning model using radiomics and clinical-radiologic features with an area under the curve (AUC) of 0.844 in the training cohort and 0.790 in the validation cohort[1]. While the authors provide clear definitions for recurrence-free survival and overall survival, it remains insufficiently detailed how ER was operationalized in terms of a specific time horizon and handling of censoring. Prognostic modeling guidelines such as TRIPOD emphasize that the time point (for example, recurrence within 1 year or 2 years) and the handling of patients without events at that horizon should be explicitly reported, because binary “ever/never” outcomes can obscure differences in follow-up and censoring[2,3]. Time-to-event models - such as Cox or flexible survival models - are widely used in contemporary radiomics studies of HCC recurrence[4], as they enable faithful modeling of recurrence-free survival, appropriate handling of censored data, and estimation of clinically relevant probabilities at fixed time horizons (e.g., 1- or 2-year recurrence risk). Clarifying the ER endpoint at a defined time horizon and aligning the analysis with established survival-modeling frameworks would therefore more accurately reflect postoperative clinical decision-making and facilitate direct comparison with existing radiomics-based prognostic tools.
Second, the radiomics workflow in the study is a major strength but would benefit from additional detail to facilitate reproducibility. Qian et al[1] obtained multiphase contrast-enhanced CT on a single General Electric Discovery HD 750 scanner with uniform acquisition parameters, and performed manual three-dimensional segmentation of tumor and liver volumes of interest using ITK-SNAP by two blinded readers with assessment of inter- and intra-observer intraclass correlation coefficients (ICCs)[1]. Radiomic features were then extracted using PyRadiomics (v3.1.0), and the authors state that standardized calculations were used[1]. This is consistent with current best practice and likely implies substantial alignment with the Image Biomarker Standardization Initiative (IBSI) framework, which has standardized definitions and reference values for 169 radiomics features and provides detailed recommendations for preprocessing, feature naming, and reporting[5]. However, full IBSI compliance requires clear reporting of all preprocessing parameters, including voxel resampling (typically to isotropic spacing), interpolation method, gray-level discretization (e.g., fixed bin width), intensity normalization, and any applied filters. Without these specifications, radiomic signatures cannot be reliably reproduced. IBSI and subsequent methodological reviews also recommend defining reproducible features using ICC thresholds - commonly ≥ 0.75-0.80[5-7]. Providing these preprocessing details and ICC criteria, ideally in supplementary material or a code repository, would enable external groups to reconstruct the feature space and compare the model against other toolkits, which is especially important given documented discrepancies among software implementations even when nominally IBSI-aligned[5-7]. Full adherence to the IBSI reporting checklist is increasingly regarded as a minimum standard for radiomics research, and a 2024 systematic review on radiomics reproducibility identified incomplete preprocessing documentation as a primary contributor to irreproducibility[8].
Third, the concept and implementation of delta-radiomics in this work is innovative but may warrant additional clarification. The authors describe delta-radiomics as using differences between arterial and portal-phase radiomics features for both tumor and background liver (“delta-T” and “delta-L”), drawing an analogy to longitudinal delta-radiomics that captures temporal changes in tumor characteristics[1]. In much of the prior literature, however, “delta-radiomics” has typically referred to changes between baseline and post-treatment or interval studies, such as pre- and post-neoadjuvant therapy in non-small cell lung cancer or longitudinal magnetic resonance imaging in HCC[9,10]. Using the same term for intra-examination phase differences may create confusion when comparing studies or synthesizing evidence. It would therefore be helpful to specify that the current work applies an “intra-examination multiphase delta-radiomics” approach and to contrast this with true longitudinal delta-radiomics, which captures temporal biological or treatment-induced changes, and aligns with terminology suggestions from a recent methodological review on temporal radiomics phenotypes[11]. This clarification also highlights the distinct biological underpinnings of the two concepts: Intra-examination arterial-portal differences mainly reflect hemodynamic and perfusion heterogeneity in cirrhotic livers, whereas longitudinal delta-radiomics reflects evolving tumor biology over time. Explicitly defining this distinction would improve methodological clarity and facilitate cross-study comparability, particularly in future meta-analyses and systematic reviews.
Fourth, Qian et al[1] combined six feature groups (arterial tumor, arterial liver, portal tumor, portal liver, tumor delta, liver delta) and used a two-step feature selection pipeline (least absolute shrinkage and selection operator and recursive feature elimination) followed by comparison of five machine-learning classifiers with five-fold cross-validation to derive the final model[1]. While such an approach is common in radiomics studies, it may introduce a risk of overfitting and selection-induced optimism. Current prediction-model reporting guidance, including TRIPOD + AI, highlights the importance of avoiding “double dipping” whereby the test or validation set influences feature selection or model tuning[3]. Nested cross-validation or bootstrap-based internal validation, in which feature selection and model tuning are repeated within each resampling loop, can provide more robust estimates of model performance and help mitigate selection bias. In this context, additional metrics such as optimism-corrected AUCs, calibration slopes, and Brier scores could complement the reported validation AUC of 0.790 and help readers assess whether the model might be over-fitted to the derivation cohort[4]. The risk of optimistic performance is substantiated by a 2024 meta-analysis of HCC recurrence radiomics, which found that studies using hold-out validation without accounting for feature-selection bias reported significantly higher performance estimates[12].
Fifth, external validity and technical generalizability deserve emphasis. The authors’ use of a single scanner and institution reduces acquisition heterogeneity and thereby strengthens internal consistency, but it also limits inference to other platforms and practice settings. Multicentre CT-based radiomics studies in HCC have shown that acquisition and reconstruction differences can affect feature stability and may degrade model performance when applied across scanners and institutions[13]. Meta-analyses of radiomics for ER prediction in HCC - including both CT and magnetic resonance imaging-based models - report encouraging pooled performance (summary AUCs around 0.85-0.90) but also highlight substantial between-study heterogeneity, modest RQS, and a scarcity of external validation[12]. In this context, an important next step would be to test the combined radiomics-clinical model developed by Qian et al[1] in an external cohort, ideally including different scanners, vendors, and patient populations (for example, varying etiologies of cirrhosis). Even a temporal validation using a later patient cohort within the same center could provide reassurance that the model’s performance remains stable over time and under evolving clinical practice. The lack of external validation remains a universal challenge that must be addressed to ensure generalizability across institutions and imaging platforms.
Sixth, several aspects of feature biology and spatial sampling could be further elaborated to enhance interpretability. Qian et al[1] show that gamma-glutamyl transferase, tumor capsule, and peritumoral enhancement emerged as independent clinical-radiologic predictors, consistent with prior literature linking these markers to microvascular invasion, aggressive tumor biology, and ER[14-16]. To improve interpretability, an explicit explanation of these associations may be helpful: Peritumoral arterial-phase enhancement has been correlated with perivascular infiltration and microvascular invasion on histopathology; incomplete or disrupted tumor capsule reflects infiltrative growth patterns; and elevated gamma-glutamyl transferase has been repeatedly associated with biologically aggressive HCC, including poor differentiation and ER. These imaging–pathology correlations suggest that the radiomic signature may be capturing underlying architectural and vascular alterations relevant to tumor aggressiveness. Meanwhile, other studies have highlighted that radiomics signatures incorporating peritumoral regions or multiple concentric rings around the tumor may improve prediction of ER and microvascular invasion beyond intertumoral features alone[17]. The current work already samples both tumor and liver volumes of interest, but the authors might consider exploring whether more explicit peritumoral shells (for example, 3-10 mm expansions) carry incremental prognostic information, as has been reported in solitary small HCC[17]. Additionally, presenting the most influential radiomics features with brief explanations (for example, whether they predominantly reflect heterogeneity, edge sharpness, or shape) and, where possible, relating them to known histopathologic correlates would facilitate clinical acceptance and aid biological interpretation. Such explainability is increasingly emphasized in contemporary radiomics guidance, including the recently proposed RQS 2.0, which introduces 'biological/clinical validation’ as a key domain[18].
Seventh, the authors appropriately assessed calibration and decision-curve analysis, and demonstrated that their combined model provides greater net benefit over a range of threshold probabilities for ER[1]. To further strengthen the clinical translation, it would be helpful to: (1) Report numerical calibration measures (intercept and slope) in the validation cohort; (2) Explore whether the optimal risk threshold identified by Youden’s index aligns with clinically acceptable trade-offs between missed ER and unnecessary intensified surveillance; and (3) Present sensitivity, specificity, and predictive values at one or more clinically meaningful cut-offs. In HCC, ER within 1-2 years is often associated with aggressive tumor biology, and some authors have advocated favoring sensitivity to avoid missing high-risk patients who might benefit from intensified surveillance or adjuvant strategies[19]. Linking these threshold decisions to concrete clinical actions - for example, 3-month vs 6-month surveillance intervals - would substantially improve clinical usability in the context of real-world decision-making.
Finally, I would like to place this work within the broader trajectory of radiomics research quality and reporting standards. Seminal contributions by Lambin et al[18] have highlighted both the promise of radiomics and the challenges of reproducibility, motivating tools such as the RQS and, more recently, RQS 2.0 and the concept of “radiomics readiness levels” to structure the pathway from discovery to clinical implementation. Systematic reviews that applied the original RQS have shown that most radiomics studies scored below 50% of the maximum, with recurrent issues including retrospective single-center designs, limited validation, incomplete reporting of imaging and feature-extraction parameters, and lack of open data or code[6]. Qian et al[1] already fulfil several higher-level criteria - such as assessing inter- and intra-observer segmentation reproducibility, performing internal validation, and using decision-curve analysis - but adopting TRIPOD + AI and RQS 2.0 explicitly (for example, by providing completed checklists as supplementary material and making the feature-extraction and modeling code publicly available where feasible) would further enhance transparency and facilitate independent validation and extension of their model[3].
In conclusion, Qian et al[1] have contributed an important and timely study on CT-based radiomics for ER prediction in HCC patients with cirrhosis, demonstrating that a combined radiomics-clinical model can effectively stratify recurrence risk and overall survival after hepatectomy. By clarifying the outcome definition and modeling framework, expanding reporting of the radiomics pipeline in line with IBSI and TRIPOD + AI, exploring the biological interpretation and spatial sampling of key features, and pursuing external validation, the authors could further enhance the interpretability, reproducibility, and generalizability of their approach. Addressing these considerations would not only strengthen the study but also elevate the model’s “radiomics readiness level”, systematically progressing it toward robust clinical utility. I congratulate the authors on their work and hope these comments are received in the constructive spirit intended, with the shared goal of advancing robust, clinically useful radiomics tools for postoperative HCC management.