TY - GEN
T1 - A Scalable External Memory Access and On-Chip Storage Architecture for Edge-AI Accelerators
T2 - 30th IEEE/ACM International Symposium on Low Power Electronics and Design, ISLPED 2025
AU - Cheng, Quan
AU - Zhang, Huizi
AU - Li, Qiufeng
AU - Liang, Yuan
AU - Zhang, Mingtao
AU - Chen, Zhenzhe
AU - Zhang, Ruilin
AU - Xiong, Jinjun
AU - Huang, Mingqiang
AU - Lin, Longyang
AU - Hashimoto, Masanori
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - For resource-constrained AI accelerators applied in edge computing, achieving high power efficiency in neural network (NN) model computation is crucial. However, current designs often overlook the efficiency of off-chip/on-chip data interaction, leading to high latency, which in turn results in suboptimal power efficiency during computation. Additionally, inefficient memory bank allocation further exacerbates latency by causing underutilization of storage resources, thereby contributing to higher overall latency and energy consumption. To address these challenges, this paper proposes a scalable multi-path rolling data refresh and layer-wise bank allocation architecture. The rolling data refresh mechanism enables efficient data interaction between off-chip and on-chip storage, reducing latency and minimizing the area overhead of on-chip memories. The layer-wise bank allocation optimizes on-chip memory utilization according to specific application requirements, improving memory efficiency. A case study on a 28nm AI accelerator demonstrates a 30.6% reduction in area, achieves a power efficiency of 7.36-10.28 TOPS/W, and reduces external memory access by 2.63% to 37.24% on VGG16 and ViT-Small.
AB - For resource-constrained AI accelerators applied in edge computing, achieving high power efficiency in neural network (NN) model computation is crucial. However, current designs often overlook the efficiency of off-chip/on-chip data interaction, leading to high latency, which in turn results in suboptimal power efficiency during computation. Additionally, inefficient memory bank allocation further exacerbates latency by causing underutilization of storage resources, thereby contributing to higher overall latency and energy consumption. To address these challenges, this paper proposes a scalable multi-path rolling data refresh and layer-wise bank allocation architecture. The rolling data refresh mechanism enables efficient data interaction between off-chip and on-chip storage, reducing latency and minimizing the area overhead of on-chip memories. The layer-wise bank allocation optimizes on-chip memory utilization according to specific application requirements, improving memory efficiency. A case study on a 28nm AI accelerator demonstrates a 30.6% reduction in area, achieves a power efficiency of 7.36-10.28 TOPS/W, and reduces external memory access by 2.63% to 37.24% on VGG16 and ViT-Small.
KW - AI accelerator
KW - edge computing
KW - layer-wise bank allocation
KW - power efficiency
KW - rolling data refresh
UR - https://www.scopus.com/pages/publications/105030043477
U2 - 10.1109/ISLPED65674.2025.11261804
DO - 10.1109/ISLPED65674.2025.11261804
M3 - Conference contribution
AN - SCOPUS:105030043477
T3 - Proceedings of the International Symposium on Low Power Electronics and Design
BT - Proceedings of the 30th International Symposium on Low Power Electronics and Design, ISLPED 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 August 2025 through 8 August 2025
ER -