TY - JOUR
T1 - Generative artificial intelligence for patient education material on gastric cancer prevention
AU - Rizkala, Tommy
AU - Muench, Natasha
AU - Hassan, Cesare
AU - Dinis-Ribeiro, Mario
AU - Tziatzios, Georgios
AU - Pimentel-Nunes, Pedro
AU - Link, Alexander
AU - Areia, Miguel
AU - Romanczyk, Marcin
AU - Matysiak-Budnik, Tamara
AU - Fernández-Esparrach, Gloria
AU - Marcos, Pedro
AU - Esposito, Gianluca
AU - Moreira, Leticia
AU - Tacheci, Ilja
AU - Santos-Antunes, João
AU - Triantafyllou, Konstantinos
AU - Bisschops, Raf
AU - Feakins, Roger
AU - Libanio, Diogo
AU - Carneiro, Fatima
AU - Bornschein, Jan
AU - Di Stefano, Luca
AU - Pugliese, Nicola
AU - Schünemann, Holger
AU - Spinelli, Antonino
AU - Repici, Alessandro
N1 - Publisher Copyright:
© 2026. Thieme. All rights reserved.
PY - 2026/6/1
Y1 - 2026/6/1
N2 - Background: This study assessed the effectiveness of large language models (LLMs) in generating lay summaries for patient education on the management of precancerous lesions and early neoplasia in the stomach. Methods: In this pilot study, we used a two-period, crossover, blinded design to compare a ChatGPT-4o summary versus a Digestive Cancers Europe (DiCE) summary. Two panels rated the materials: expert physicians and DiCE Patient Advisory Committee members. Experts scored accuracy, completeness, comprehensibility, and satisfaction (across five sections); patients rated overall completeness, comprehensibility, and satisfaction. Paired comparisons used mixed-effects estimates. Readability was assessed with Flesch-Kincaid grade level (FKGL) and SMOG index. Results: Median expert ratings were similar between materials across metrics. For the overall summary, median (range; IQR) scores were: accuracy 5 (4-6; 1) for ChatGPT-4o vs. 5 (3-6; 1) for DiCE (P=0.10); completeness 4 (3-5; 1) vs. 4 (2-5; 1; P=0.27); comprehensibility 4 (3-5; 1) vs. 4 (2-5; 1; P=0.33); and satisfaction 4 (2-5; 1) vs. 3 (1-5; 2; P=0.53). Patient ratings mirrored experts, with very similar results. Readability failed to meet guideline recommendations for both summaries on both FKGL and SMOG scores. Conclusion: ChatGPT-4o produced patient materials comparable to DiCE, but both require readability optimization; a human-in-the-loop workflow and future tests across prompts and models are warranted.
AB - Background: This study assessed the effectiveness of large language models (LLMs) in generating lay summaries for patient education on the management of precancerous lesions and early neoplasia in the stomach. Methods: In this pilot study, we used a two-period, crossover, blinded design to compare a ChatGPT-4o summary versus a Digestive Cancers Europe (DiCE) summary. Two panels rated the materials: expert physicians and DiCE Patient Advisory Committee members. Experts scored accuracy, completeness, comprehensibility, and satisfaction (across five sections); patients rated overall completeness, comprehensibility, and satisfaction. Paired comparisons used mixed-effects estimates. Readability was assessed with Flesch-Kincaid grade level (FKGL) and SMOG index. Results: Median expert ratings were similar between materials across metrics. For the overall summary, median (range; IQR) scores were: accuracy 5 (4-6; 1) for ChatGPT-4o vs. 5 (3-6; 1) for DiCE (P=0.10); completeness 4 (3-5; 1) vs. 4 (2-5; 1; P=0.27); comprehensibility 4 (3-5; 1) vs. 4 (2-5; 1; P=0.33); and satisfaction 4 (2-5; 1) vs. 3 (1-5; 2; P=0.53). Patient ratings mirrored experts, with very similar results. Readability failed to meet guideline recommendations for both summaries on both FKGL and SMOG scores. Conclusion: ChatGPT-4o produced patient materials comparable to DiCE, but both require readability optimization; a human-in-the-loop workflow and future tests across prompts and models are warranted.
UR - https://www.scopus.com/pages/publications/105030351666
U2 - 10.1055/a-2780-0664
DO - 10.1055/a-2780-0664
M3 - Article
C2 - 41688051
AN - SCOPUS:105030351666
SN - 0013-726X
VL - 58
SP - 669
EP - 677
JO - Endoscopy
JF - Endoscopy
IS - 6
ER -