Published online Oct 7, 2026. doi: 10.3748/wjg.119857
Revised: March 20, 2026
Accepted: May 28, 2026
Published online: October 7, 2026
Processing time: 205 Days and 19.7 Hours
Large language models (LLMs) are increasingly used for patient education, but their reliability for Chinese Helicobacter pylori (H. pylori)-related counseling remains unclear.
To evaluate the performance of LLMs in Chinese H. pylori-related question answering across structured tests, guideline-based questions, and real-world patient queries.
This three-phase comparative study evaluated four LLMs. In phase 1, models answered 112 H. pylori single-best-answer questions. In phase 2, they answered 30 guideline-based clinical questions, and three blinded senior gastroenterologists rated correctness, completeness, readability, helpfulness, and safety. Readability was also assessed using the language difficulty understanding tool for general public. Based on performance in phases 1 and 2, two models were selected for phase 3, in which 40 patients provided 120 real-world questions. Clinicians rated responses, and patients rated satisfaction and per
In phase 1, accuracy ranged from 75.0% to 85.0%. In phase 2, ChatGPT5 achieved the highest scores across all domains, including correctness (4.63), completeness (4.89), readability (4.88), helpfulness (4.86), and safety (4.59). In phase 3, DeepSeek outperformed ChatGPT5 in correctness (4.80 vs 3.85, P < 0.001), completeness (4.80 vs 4.05, P = 0.008), perceived readability (4.65 vs 4.13, P = 0.006), and patient satisfaction (4.73 vs 4.05, P < 0.001).
LLM performance in H. pylori-related health-information support was task dependent. Strong performance in structured assessments did not necessarily predict better real-world patient interaction. Multidimensional eva
Core Tip: This three-phase study evaluated large language models for Chinese Helicobacter pylori health-information support, progressing from multiple-choice testing to guideline-based questions and real-world patient queries. Performance was clearly task dependent: The model that performed best in structured assessments was not necessarily the most effective in patient interaction. In phase 3, DeepSeek showed better correctness, completeness, perceived readability, and patient satisfaction. These findings support multidimensional evaluation and clinician oversight before patient-facing deployment.