TY - GEN
T1 - Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines - A Comparative Analysis
AU - Kumar, Punit
AU - Imran, Asif
AU - Kosar, Tevfik
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/7/6
Y1 - 2026/7/6
N2 - This paper presents a comparative performance analysis of three popular Python data manipulation libraries - Pandas, Polars, and Dask - within the context of deep learning training pipelines. The existing studies in this area do not embed the libraries inside a full deep-learning training pipeline where data loading, preprocessing, and batch feeding interact tightly with GPU workloads. To bridge this gap, we integrate Pandas, Polars, and Dask into representative deep learning training and inference pipelines and conduct experiments across a wide range of various machine learning models and datasets, measuring key performance indicators such as runtime, memory usage, disk usage, and energy consumption (CPU and GPU). Our comprehensive analysis reveals that Polars consistently minimizes CPU energy consumption on larger workloads, while Pandas remains competitive for moderate sizes. Dask's overhead can lead to higher energy usage on small to moderate datasets. All three libraries achieve similar runtimes for heavy GPU workloads (ResNet, Mask R-CNN). Polars and Pandas maintain lower CPU memory footprints than Dask, but Dask offers easier scalability if data truly exceeds available RAM. Polars shows marginal energy savings on the CPU during preprocessing.
AB - This paper presents a comparative performance analysis of three popular Python data manipulation libraries - Pandas, Polars, and Dask - within the context of deep learning training pipelines. The existing studies in this area do not embed the libraries inside a full deep-learning training pipeline where data loading, preprocessing, and batch feeding interact tightly with GPU workloads. To bridge this gap, we integrate Pandas, Polars, and Dask into representative deep learning training and inference pipelines and conduct experiments across a wide range of various machine learning models and datasets, measuring key performance indicators such as runtime, memory usage, disk usage, and energy consumption (CPU and GPU). Our comprehensive analysis reveals that Polars consistently minimizes CPU energy consumption on larger workloads, while Pandas remains competitive for moderate sizes. Dask's overhead can lead to higher energy usage on small to moderate datasets. All three libraries achieve similar runtimes for heavy GPU workloads (ResNet, Mask R-CNN). Polars and Pandas maintain lower CPU memory footprints than Dask, but Dask offers easier scalability if data truly exceeds available RAM. Polars shows marginal energy savings on the CPU during preprocessing.
KW - Dask
KW - Deep Learning
KW - End-to-end Pipeline
KW - Energy Efficiency
KW - Inference
KW - Pandas
KW - Performance Evaluation
KW - Polars
KW - Training
UR - https://www.scopus.com/pages/publications/105044692504
U2 - 10.1145/3786148.3788619
DO - 10.1145/3786148.3788619
M3 - Conference contribution
AN - SCOPUS:105044692504
T3 - Proceedings - 2026 IEEE/ACM 10th International Workshop on Green and Sustainable Software, GREENS 2026
SP - 14
EP - 21
BT - Proceedings - 2026 IEEE/ACM 10th International Workshop on Green and Sustainable Software, GREENS 2026
PB - Association for Computing Machinery, Inc
T2 - 10th International Workshop on Green and Sustainable Software, GREENS 2026
Y2 - 12 April 2026 through 18 April 2026
ER -