Abstract
The CellCards knowledgebase aims to systematically gather, and represent individual cell types. This study presents our development of a dynamic extraction, transformation, and loading (ETL) pipeline designed to automatically populate the CellCards database with a vast array of cells from ontologies, including the Cell Ontology (CL) and Cell Line Ontology (CLO). The CellCards database schema includes five tables, with a key feature being the use of one table to encompass all necessary terms from the ontologies and another table to outline the relationships among these terms. The ETL process is powered by a Python script that embeds SPARQL queries directed at the Ontobee SPARQL endpoint. The final ETL program successfully extracted and loaded over 3,500 cell types from CL and 40,000 cell line entries from CLO into the new CellCards database, including the cell type name, parent cell type, synonyms, anatomical locations, etc. The gene biomarkers of cells were automatically extracted from the Common Coordinate Framework Ontology (CCFO). This enhanced database will be used to update the website and query program, with these updates scheduled for summer 2024.
| Original language | English |
|---|---|
| Journal | CEUR Workshop Proceedings |
| Volume | 3939 |
| State | Published - 2024 |
| Event | 15th International Conference on Biological and Biomedical Ontology, ICBO 2024 - Enschede, Netherlands Duration: Jul 18 2024 → Jul 19 2024 |
Keywords
- Cell
- Cell Line Ontology
- Cell Ontology
- Common Coordinate Framework Ontology
- ETL (Extract, Transform, Load)
- Gene Ontology
- Knowledgebase
- Multi-Threading
- Ontology
- Python
- SPARQL
- Uber-anatomy Ontology
Fingerprint
Dive into the research topics of 'CellCards: Development of a dynamic ontology-derived ETL pipeline for automatic cell information extraction and analysis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver