Skip to main navigation Skip to search Skip to main content

Improving Joint Learning of Chest X-Ray and Radiology Report by Word Region Alignment

  • Zhanghexuan Ji
  • , Mohammad Abuzar Shaikh
  • , Dana Moukheiber
  • , Sargur N. Srihari
  • , Yifan Peng
  • , Mingchen Gao
  • SUNY Buffalo
  • Cornell University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

13 Scopus citations

Abstract

Self-supervised learning provides an opportunity to explore unlabeled chest X-rays and their associated free-text reports accumulated in clinical routine without manual supervision. This paper proposes a Joint Image Text Representation Learning Network (JoImTeRNet) for pre-training on chest X-ray images and their radiology reports. The model was pre-trained on both the global image-sentence level and the local image region-word level for visual-textual matching. Both are bidirectionally constrained on Cross-Entropy based and ranking-based Triplet Matching Losses. The region-word matching is calculated using the attention mechanism without direct supervision about their mapping. The pre-trained multi-modal representation learning paves the way for downstream tasks concerning image and/or text encoding. We demonstrate the representation learning quality by cross-modality retrievals and multi-label classifications on two datasets: OpenI-IU and MIMIC-CXR. Our code is available at https://github.com/mshaikh2/JoImTeR_MLMI_2021.

Original languageEnglish
Title of host publicationMachine Learning in Medical Imaging - 12th International Workshop, MLMI 2021, Held in Conjunction with MICCAI 2021, Proceedings
EditorsChunfeng Lian, Xiaohuan Cao, Islem Rekik, Xuanang Xu, Pingkun Yan
PublisherSpringer Science and Business Media Deutschland GmbH
Pages110-119
Number of pages10
ISBN (Print)9783030875886
DOIs
StatePublished - 2021
Event12th International Workshop on Machine Learning in Medical Imaging, MLMI 2021, held in conjunction with 24th International Conference on Medical Image Computing and Computer Assisted Intervention, MICCAI 2021 - Virtual, Online
Duration: Sep 27 2021Sep 27 2021

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume12966 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference12th International Workshop on Machine Learning in Medical Imaging, MLMI 2021, held in conjunction with 24th International Conference on Medical Image Computing and Computer Assisted Intervention, MICCAI 2021
CityVirtual, Online
Period09/27/2109/27/21

Keywords

  • Attention
  • Multi-modality
  • Self-supervised learning

Fingerprint

Dive into the research topics of 'Improving Joint Learning of Chest X-Ray and Radiology Report by Word Region Alignment'. Together they form a unique fingerprint.

Cite this