Skip to main navigation Skip to search Skip to main content

Use of artificial intelligence chatbots in clinical management of immune-related adverse events

  • Hannah Burnette
  • , Aliyah Pabani
  • , Mitchell S. Von Itzstein
  • , Benjamin Switzer
  • , Run Fan
  • , Fei Ye
  • , Igor Puzanov
  • , Jarushka Naidoo
  • , Paolo A. Ascierto
  • , David E. Gerber
  • , Marc S. Ernstoff
  • , Douglas B. Johnson
  • Vanderbilt University
  • Johns Hopkins University
  • University of Texas Southwestern Medical Center
  • Roswell Park Cancer Institute
  • Royal College of Surgeons in Ireland
  • IRCCS Istituto nazionale tumori Fondazione Giovanni Pascale - Napoli
  • National Institutes of Health

Research output: Contribution to journalArticlepeer-review

16 Scopus citations

Abstract

Background Artificial intelligence (AI) chatbots have become a major source of general and medical information, though their accuracy and completeness are still being assessed. Their utility to answer questions surrounding immune-related adverse events (irAEs), common and potentially dangerous toxicities from cancer immunotherapy, are not well defined. Methods We developed 50 distinct questions with answers in available guidelines surrounding 10 irAE categories and queried two AI chatbots (ChatGPT and Bard), along with an additional 20 patient-specific scenarios. Experts in irAE management scored answers for accuracy and completion using a Likert scale ranging from 1 (least accurate/complete) to 4 (most accurate/complete). Answers across categories and across engines were compared. Results Overall, both engines scored highly for accuracy (mean scores for ChatGPT and Bard were 3.87 vs 3.5, p<0.01) and completeness (3.83 vs 3.46, p<0.01). Scores of 1-2 (completely or mostly inaccurate or incomplete) were particularly rare for ChatGPT (6/800 answer-ratings, 0.75%). Of the 50 questions, all eight physician raters gave ChatGPT a rating of 4 (fully accurate or complete) for 22 questions (for accuracy) and 16 questions (for completeness). In the 20 patient scenarios, the average accuracy score was 3.725 (median 4) and the average completeness was 3.61 (median 4). Conclusions AI chatbots provided largely accurate and complete information regarding irAEs, and wildly inaccurate information ("hallucinations") was uncommon.

Original languageEnglish
Article numbere008599
JournalJournal for ImmunoTherapy of Cancer
Volume12
Issue number5
DOIs
StatePublished - May 30 2024

Keywords

  • Colitis
  • Immune Checkpoint Inhibitor
  • Immune related adverse event - irAE
  • Pneumonitis
  • Thyroiditis

Fingerprint

Dive into the research topics of 'Use of artificial intelligence chatbots in clinical management of immune-related adverse events'. Together they form a unique fingerprint.

Cite this