Skip to main navigation Skip to search Skip to main content

Generative artificial intelligence for patient education material on gastric cancer prevention

  • Tommy Rizkala
  • , Natasha Muench
  • , Cesare Hassan
  • , Mario Dinis-Ribeiro
  • , Georgios Tziatzios
  • , Pedro Pimentel-Nunes
  • , Alexander Link
  • , Miguel Areia
  • , Marcin Romanczyk
  • , Tamara Matysiak-Budnik
  • , Gloria Fernández-Esparrach
  • , Pedro Marcos
  • , Gianluca Esposito
  • , Leticia Moreira
  • , Ilja Tacheci
  • , João Santos-Antunes
  • , Konstantinos Triantafyllou
  • , Raf Bisschops
  • , Roger Feakins
  • , Diogo Libanio
  • Fatima Carneiro, Jan Bornschein, Luca Di Stefano, Nicola Pugliese, Holger Schünemann, Antonino Spinelli, Alessandro Repici
  • IRCCS Istituto Clinico Humanitas - Rozzano (Milano)
  • DiCE: Digestive Cancer Europe
  • Humanitas University
  • Instituto Português de Oncologia do Porto Francisco Gentil E.P.E.
  • National and Kapodistrian University of Athens
  • University of Porto
  • Otto von Guericke University Magdeburg
  • Portuguese Oncology Institute
  • Academy of Silesia
  • Endoterapia
  • Nantes University Hospital
  • Nantes Université
  • University of Barcelona
  • August Pi i Sunyer Biomedical Research Institute
  • Centro de Investigación Biomédica en Red de Enfermedades Hepáticas y Digestivas
  • Pêro da Covilhã Hospital
  • University of Beira Interior
  • University of Rome "sapienza,"
  • Charles University
  • KU Leuven
  • Royal Free London NHS Foundation Trust
  • University College London
  • Faculty of Medicine
  • University of Oxford

Research output: Contribution to journalArticlepeer-review

Abstract

Background: This study assessed the effectiveness of large language models (LLMs) in generating lay summaries for patient education on the management of precancerous lesions and early neoplasia in the stomach. Methods: In this pilot study, we used a two-period, crossover, blinded design to compare a ChatGPT-4o summary versus a Digestive Cancers Europe (DiCE) summary. Two panels rated the materials: expert physicians and DiCE Patient Advisory Committee members. Experts scored accuracy, completeness, comprehensibility, and satisfaction (across five sections); patients rated overall completeness, comprehensibility, and satisfaction. Paired comparisons used mixed-effects estimates. Readability was assessed with Flesch-Kincaid grade level (FKGL) and SMOG index. Results: Median expert ratings were similar between materials across metrics. For the overall summary, median (range; IQR) scores were: accuracy 5 (4-6; 1) for ChatGPT-4o vs. 5 (3-6; 1) for DiCE (P=0.10); completeness 4 (3-5; 1) vs. 4 (2-5; 1; P=0.27); comprehensibility 4 (3-5; 1) vs. 4 (2-5; 1; P=0.33); and satisfaction 4 (2-5; 1) vs. 3 (1-5; 2; P=0.53). Patient ratings mirrored experts, with very similar results. Readability failed to meet guideline recommendations for both summaries on both FKGL and SMOG scores. Conclusion: ChatGPT-4o produced patient materials comparable to DiCE, but both require readability optimization; a human-in-the-loop workflow and future tests across prompts and models are warranted.

Original languageEnglish
Pages (from-to)669-677
Number of pages9
JournalEndoscopy
Volume58
Issue number6
DOIs
StatePublished - Jun 1 2026

Fingerprint

Dive into the research topics of 'Generative artificial intelligence for patient education material on gastric cancer prevention'. Together they form a unique fingerprint.

Cite this