BPG is committed to discovery and dissemination of knowledge
Observational Study
Copyright: ©Author(s) 2026.
World J Gastroenterol. Oct 7, 2026; 32(37): 119857
Published online Oct 7, 2026. doi: 10.3748/wjg.119857
Figure 1
Figure 1 Overall study design and workflow of the three-phase evaluation framework. MCQ: Multiple-choice question; LLM: Large language model; LDU-TGP: Ludong University Text Grading Platform.
Figure 2
Figure 2 Performance comparison of large language models on multiple-choice questions. H. pylori: Helicobacter pylori; MCQ: Multiple-choice question.
Figure 3
Figure 3 Evaluation results of large language models on clinical questions. aP < 0.05, bP < 0.01, cP < 0.001, and dP < 0.00001.
Figure 4
Figure 4  Multidimensional assessment of large language models on patient-generated questions.
Figure 5
Figure 5 Radar chart summarizing the performance of different models across multiple evaluation dimensions. aP < 0.05, bP < 0.01.


Write to the Help Desk