REVISTA
Algorithms

TODAS

Inicio / Algorithms / Vol: 16 Par: 2 (2023) / Art�culo

ART�CULO

TITULO

Model Parallelism Optimization for CNN FPGA Accelerator

Jinnan Wang

Weiqin Tong and Xiaoli Zhi

Resumen

Convolutional neural networks (CNNs) have made impressive achievements in image classification and object detection. For hardware with limited resources, it is not easy to achieve CNN inference with a large number of parameters without external storage. Model parallelism is an effective way to reduce resource usage by distributing CNN inference among several devices. However, parallelizing a CNN model is not easy, because CNN models have an essentially tightly-coupled structure. In this work, we propose a novel model parallelism method to decouple the CNN structure with group convolution and a new channel shuffle procedure. Our method could eliminate inter-device synchronization while reducing the memory footprint of each device. Using the proposed model parallelism method, we designed a parallel FPGA accelerator for the classic CNN model ShuffleNet. This accelerator was further optimized with features such as aggregate read and kernel vectorization to fully exploit the hardware-level parallelism of the FPGA. We conducted experiments with ShuffleNet on two FPGA boards, each of which had an Intel Arria 10 GX1150 and 16GB DDR3 memory. The experimental results showed that when using two devices, ShuffleNet achieved a 1.42� speed increase and reduced its memory footprint by 34%, as compared to its non-parallel counterpart, while maintaining accuracy.

Palabras claves

convolution neural network - model parallelism - field programmable gate array - inference accelerating

Acceso

P�GINAS

pp. 0 - 0

N�MERO

Volumen: 16 Parte: 2 (2023)

MATERIAS

INGENIER�A Y CONSTRUCCI�N CIVIL
TECNOLOG�A

REVISTAS SIMILARES

Algorithms
Applied Sciences
Water

DOI

https://doi.org/10.3390/a16020110

Art�culos similares

An Architecture for a Tri-Programming Model-Based Parallel Hybrid Testing Tool

Acceso

Saeed Musaad Altalhi, Fathy Elbouraey Eassa, Abdullah Saad Al-Malaise Al-Ghamdi, Sanaa Abdullah Sharaf, Ahmed Mohammed Alghamdi, Khalid Ali Almarhabi and Maher Ali Khemakhem

As the development of high-performance computing (HPC) is growing, exascale computing is on the horizon. Therefore, it is imperative to develop parallel systems, such as graphics processing units (GPUs) and programming models, that can effectively utilis... ver m�s

Revista: Applied Sciences

P System with Fractional Reduction

Acceso

Hai Nan, Yumeng Kong, Jie Zhan, Mingqiang Zhou and Ling Bai

Membrane computing is a branch of natural computing, which is a new computational model abstracted from the study of the function and structure of living biological cells. The study of numerical computation based on membrane computation has received incr... ver m�s

Revista: Applied Sciences

Performance Estimation of High-Level Dataflow Program on Heterogeneous Platforms by Dynamic Network Execution

Acceso

Aurelien Bloch, Simone Casale-Brunet and Marco Mattavelli

The performance of programs executed on heterogeneous parallel platforms largely depends on the design choices regarding how to partition the processing on the various different processing units. In other words, it depends on the assumptions and paramete... ver m�s

Revista: Journal of Low Power Electronics and Applications

A Neurally Inspired Model of Figure Ground Organization with Local and Global Cues

Acceso

Sudarshan Ramenahalli

Figure Ground Organization (FGO)-inferring spatial depth ordering of objects in a visual scene-involves determining which side of an occlusion boundary is figure (closer to the observer) and which is ground (further away from the observer). A combination... ver m�s

Revista: AI

Performance Analysis of Thread Block Schedulers in GPGPU and Its Implications

Acceso

KyungWoon Cho and Hyokyung Bahn

GPGPU (General-Purpose Graphics Processing Unit) consists of hardware resources that can execute tens of thousands of threads simultaneously. However, in reality, the parallelism is limited as resource allocation is performed by the base unit called thre... ver m�s

Revista: Applied Sciences

Revistas destacadas

Acceso directo a los n�meros publicados en la revista Infrastructures

Infrastructures

Acceso directo a los n�meros publicados en la revista Informed Infraestructure

Informed Infraestructure

Acceso directo a los n�meros publicados en la revista BiT

Acceso directo a los n�meros publicados en la revista Revista de la Construcci�n

Revista de la Construcci�n

Ver todas las revistas disponibles