Skip to main navigation Skip to search Skip to main content

Analysis of Meta-Learning Approaches for TCGA Pan-cancer Datasets

  • Jingyuan Chou
  • , Stefan Bekiranov
  • , Chongzhi Zang
  • , Mengdi Huai
  • , Aidong Zhang
  • University of Virginia

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

4 Scopus citations

Abstract

Cancer has been characterized as a heterogeneous disease, and the classification of cancer subtypes has become a necessity in cancer research, as it can facilitate the subsequent clinical management of patients and provide clinical decision support for clinicians. With the advance of machine learning in the last decade, many researchers employ machine learning to tackle the cancer classification problem. Importantly, traditional machine learning algorithms require a large amount of annotated data for model training. However, collection of large amounts of annotated data is time-consuming and expensive and may not be realistic in real-world activities. Facing data scarcity, metalearning is proposed to tackle this problem. Meta-learning utilizes prior knowledge learned from related tasks and generalizes to new tasks of limited supervised experience, and it has been applied in many fields to tackle scarce annotated data problem, such as few-shot image classification, drug discovery, etc. As data scarcity is common in cancer research and diagnosis studies, and there are only few previous studies that classify cancers based on limited annotated data. We explore the meta-learning algorithm (MAML) to tackle the scenario where only limited annotated data are available. In this work, our objective is to comprehensively compare MAML among few-shot learning methods (matching network and prototypical network) and traditional machine learning methods (random forest and Knearest neighbor). Experimental results on The Cancer Genome Atlas (TCGA) cancer patient data demonstrates the effectiveness and superiority of MAML over other methods, including its ability to outperform the other methods using 4.5-fold fewer features.

Original languageEnglish
Title of host publicationProceedings - 2020 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2020
EditorsTaesung Park, Young-Rae Cho, Xiaohua Tony Hu, Illhoi Yoo, Hyun Goo Woo, Jianxin Wang, Julio Facelli, Seungyoon Nam, Mingon Kang
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages257-262
Number of pages6
ISBN (Electronic)9781728162157
DOIs
StatePublished - Dec 16 2020
Event2020 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2020 - Virtual, Seoul, Korea, Republic of
Duration: Dec 16 2020Dec 19 2020

Publication series

NameProceedings - 2020 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2020

Conference

Conference2020 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2020
Country/TerritoryKorea, Republic of
CityVirtual, Seoul
Period12/16/2012/19/20

Keywords

  • cancer genomics
  • cancer proteomics
  • Meta-Learning
  • pan-cancer analysis

Fingerprint

Dive into the research topics of 'Analysis of Meta-Learning Approaches for TCGA Pan-cancer Datasets'. Together they form a unique fingerprint.

Cite this