Skip to main navigation Skip to search Skip to main content

Recalibrated cross-modal alignment network for radiology report generation with weakly supervised contrastive learning

  • Xiaodi Hou
  • , Xiaobo Li
  • , Zhi Liu
  • , Shengtian Sang
  • , Mingyu Lu
  • , Yijia Zhang
  • Dalian Maritime University

Research output: Contribution to journalArticlepeer-review

22 Scopus citations

Abstract

Automatic radiology report generation is rapidly becoming an essential method for medical diagnosis and precision medicine, which will help clinical doctors make more informed decisions and achieve better results. Most previous studies mainly employ an encoder–decoder architecture and tend to focus more on the text generation part, which ignore the following problems: (1) visual and textual data bias; (2) inadequate cross-modal interaction. In this paper, we propose a Recalibrated Cross-modal Alignment Network for radiology report generation with weakly supervised contrastive learning (RCAN). Specifically, we design a recalibrated visual extractor to extract critical abnormal features. To achieve effective alignment between medical images and reports, we develop a cross-modal gated memory matrix. We effectively map cross-modal information by introducing gating units to memorize interactions between different modalities selectively. In addition, to encourage more accurate image anomaly detection and recognition, we propose a new weakly supervised contrastive learning approach. We conduct extensive experimental verification on two real-world datasets, IU X-ray and MIMIC-CXR, and the results show that our model can more accurately identify abnormal features in images and generate more continuous medical reports. Importantly, the proposed RCAN achieves state-of-the-art performance compared with competitive models, particularly achieving improvements of 2.6%, 1.6%, 1.4%, 1.6%, 0.3% and 0.9%, 1.7%, 1.6%, 1.2%, 0.9% in BLEU-1, BLEU-2, BLEU-3, BLEU-4, METEOR metrics on both datasets, respectively.

Original languageEnglish
Article number126394
JournalExpert Systems with Applications
Volume269
DOIs
StatePublished - Apr 15 2025

Keywords

  • Automatic radiology report generation
  • Cross-modal alignment
  • Visual and textual data bias
  • Weakly supervised contrastive learning

Fingerprint

Dive into the research topics of 'Recalibrated cross-modal alignment network for radiology report generation with weakly supervised contrastive learning'. Together they form a unique fingerprint.

Cite this