BPG is committed to discovery and dissemination of knowledge
Observational Study
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Gastroenterol. Oct 7, 2026; 32(37): 119857
Published online Oct 7, 2026. doi: 10.3748/wjg.119857
Evaluating large language models in Helicobacter pylori-related question answering: From knowledge tests to patient queries
Shi-Ping Sun, Dan-Ye Niu, Ming-Kai Yuan, Li Liu, Yi Li, Han Min
Shi-Ping Sun, Yi Li, Han Min, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Dan-Ye Niu, Department of Clinical Nutrition, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Ming-Kai Yuan, Department of Hepatobiliary and Pancreatic Surgery, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Li Liu, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Suzhou Women and Children’s Health Hospital, Suzhou 215000, Jiangsu Province, China
Co-first authors: Shi-Ping Sun and Dan-Ye Niu.
Author contributions: Sun SP, Niu DY, and Min H conceptualized the study, and drafted the manuscript; Sun SP, Niu DY, and Yuan MK developed the methodology; Sun SP, Niu DY, Li Y, and Min H curated the data; Sun SP and Niu DY performed the formal analysis, and contributed equally as co-first authors; Sun SP, Niu DY, Li Y, and Min H conducted the investigation; Sun SP, Liu L, and Min H reviewed and edited the manuscript; Yuan MK, Liu L and Min H supervised the study. All authors approved the final version to publish.
AI contribution statement: No AI tools (including ChatGPT, Grammarly, DeepL, or any other AI-based tools) were used in the preparation of this manuscript. No part of the main text (including the Abstract, Introduction, Materials and Methods, Results, Discussion, or Conclusion) was generated by AI. AI tools were not used for language polishing, translation, data analysis, or writing assistance. AI tools did not participate in the study design, data analysis, or interpretation of results. No images in the manuscript were generated by AI. We confirm that the manuscript was entirely prepared by the authors.
Supported by Suzhou Major Disease Multicenter Clinical Research Project, No. DZXYJ202508.
Institutional review board statement: This study was approved by the Ethical Committee of the Suzhou Hospital Affiliated to Nanjing Medical University, No. K-2025-162-K01.
Informed consent statement: All participants provided informed consent prior to participation.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
STROBE statement: The authors have read the STROBE Statement-checklist of items, and the manuscript was prepared and revised according to the STROBE Statement-checklist of items.
Data sharing statement: All data are included in the article and its Supplementary material.
Corresponding author: Han Min, MD, Full Professor, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, No. 26 Daoqian Street, Gusu District, Suzhou 215000, Jiangsu Province, China. minhan1981@njmu.edu.cn
Received: February 9, 2026
Revised: March 20, 2026
Accepted: May 28, 2026
Published online: October 7, 2026
Processing time: 205 Days and 19.7 Hours
Abstract
BACKGROUND

Large language models (LLMs) are increasingly used for patient education, but their reliability for Chinese Helicobacter pylori (H. pylori)-related counseling remains unclear.

AIM

To evaluate the performance of LLMs in Chinese H. pylori-related question answering across structured tests, guideline-based questions, and real-world patient queries.

METHODS

This three-phase comparative study evaluated four LLMs. In phase 1, models answered 112 H. pylori single-best-answer questions. In phase 2, they answered 30 guideline-based clinical questions, and three blinded senior gastroenterologists rated correctness, completeness, readability, helpfulness, and safety. Readability was also assessed using the language difficulty understanding tool for general public. Based on performance in phases 1 and 2, two models were selected for phase 3, in which 40 patients provided 120 real-world questions. Clinicians rated responses, and patients rated satisfaction and perceived readability.

RESULTS

In phase 1, accuracy ranged from 75.0% to 85.0%. In phase 2, ChatGPT5 achieved the highest scores across all domains, including correctness (4.63), completeness (4.89), readability (4.88), helpfulness (4.86), and safety (4.59). In phase 3, DeepSeek outperformed ChatGPT5 in correctness (4.80 vs 3.85, P < 0.001), completeness (4.80 vs 4.05, P = 0.008), perceived readability (4.65 vs 4.13, P = 0.006), and patient satisfaction (4.73 vs 4.05, P < 0.001).

CONCLUSION

LLM performance in H. pylori-related health-information support was task dependent. Strong performance in structured assessments did not necessarily predict better real-world patient interaction. Multidimensional evaluation and clinician oversight are needed before patient-facing use.

Keywords: Helicobacter pylori; Large language models; Patient education; Readability; Patient satisfaction; Clinical decision support systems; Artificial intelligence

Core Tip: This three-phase study evaluated large language models for Chinese Helicobacter pylori health-information support, progressing from multiple-choice testing to guideline-based questions and real-world patient queries. Performance was clearly task dependent: The model that performed best in structured assessments was not necessarily the most effective in patient interaction. In phase 3, DeepSeek showed better correctness, completeness, perceived readability, and patient satisfaction. These findings support multidimensional evaluation and clinician oversight before patient-facing deployment.

Write to the Help Desk