• Home
  • Search
  • Hardware Implementation of Depthwise Separable Convolution Neural Network
  • Cite Icon5
  • https://doi.org/10.1109/icsict49897.2020.9278300Copy DOI Icon

Hardware Implementation of Depthwise Separable Convolution Neural Network

  • Nov 3, 2020
  • Yancao Jiang +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In this paper, an efficient architecture of depthwise separable convolutional neural network is presented. By simplifying the network architecture and adopting 16-bit fixed point form, the computation cost is substantially reduced with slightly decrease in image classification precision based on CIFAR10 dataset. In order to improve energy efficiency and reduce memory access, the custom processing element (PE) is proposed, which supports zero skipping, local accumulation and memory, as well as function multiplexing (for different convolution operations). Besides, the specific hardware architecture based on our custom PE is proposed and the hardware architecture can also support two dataflow modes in order to customize traditional convolution and depthwise separable convolution (DWC) dataflow. The hardware architecture is implemented on a Xilinx Zynq 7Z020 field-programmable gate array (FPGA) platform and the experimental results are implemented. By exploiting parallelism and data reuse, post synthesis simulation with a clock frequency of 50 MHz shows that the network achieves a peak performance of 9.6 GOPS and an energy efficiency of over 90.56 GOPS/W for single-frame runtime inference, achieving a 59× higher energy efficiency compared with the CPU Intel i5-8400. The results show that the proposed accelerator can classify each picture from Cifar10 in 10 ms, which is about 100 frames per second.The FPGA design achieves 41x speedup if compared to CPU, achieving real-time image classification at 100 fps.

Similar Papers
  • Research Article
  • Citations34

A classification model based on depthwise separable convolutional neural network to identify rice plant diseases

  • Aug 01, 2022
  • International Journal of Electrical and Computer Engineering (IJECE)
  • Md Sazzadul Islam Prottasha +1
  • Research Article
  • Citations19

An Efficient FPGA-based Depthwise Separable Convolutional Neural Network Accelerator with Hardware Pruning

  • Feb 12, 2024
  • ACM Transactions on Reconfigurable Technology and Systems
  • Zhengyan Liu +3
  • Conference Article

Evaluating CNNs: SqueezeNet and MobileNet on Corrupted Images

  • Aug 12, 2025
  • Noorul Julaiha Ag +5
  • Research Article
  • Citations2

Comparison of CNN Architecture of Image Classification Using CIFAR10 Datasets

  • Dec 21, 2023
  • International Journal on Engineering Technology
  • Yogesh Pant +4
  • PDF
  • Research Article
  • Citations14

Fault Line Selection Method Based on Transfer Learning Depthwise Separable Convolutional Neural Network

  • Nov 10, 2021
  • Journal of Electrical and Computer Engineering
  • Haixia Zhang +1
  • Book Chapter
  • Citations1

Self-build Deep Convolutional Neural Network Architecture Using Evolutionary Algorithms

  • Jan 01, 2023
  • Vidyanand Mishra +1
  • PDF
  • Research Article
  • Citations7

FPGA-Flux Proprietary System for Online Detection of Outer Race Faults in Bearings

  • Apr 19, 2023
  • Electronics
  • Jonathan Cureño-Osornio +5
  • Conference Article
  • Citations5

Wearable FPGA Platform for Accelerated DSP and AI Applications

  • Mar 21, 2022
  • Daniel Roggen +3
  • Conference Article
  • Citations2

An Efficient Accelerator with Winograd for Novel Convolutional Neural Networks

  • May 13, 2022
  • Zhijian Lin +3
  • Conference Article
  • Citations2

Unleashing the computational power of FPGAs to efficiently perform SPMV operation

  • Nov 15, 2021
  • Federico Favaro +2
  • Research Article
  • Citations4

Human Activity Recognition in a Realistic and Multiview Environment Based on Two-Dimensional Convolutional Neural Network

  • May 09, 2023
  • Journal of Artificial Intelligence and Technology
  • Ashish Khare +2
  • Conference Article

Hardware Design for Implementation of Human Activity Recognition Based on a U-Net++ Network

  • Oct 25, 2025
  • Ang Li +1
  • Conference Article
  • Citations18

Efficient Optimization and Hardware Acceleration of CNNs towards the Design of a Scalable Neuro inspired Architecture in Hardware

  • Jan 01, 2018
  • The H Vu +3
  • Conference Article

Evaluation of FPGA Acceleration of Neural Networks

  • Dec 11, 2023
  • Kalpa publications in computing
  • Emil Stevnsborg +3
  • Research Article
  • Citations6

Empowering edge devices: FPGA‐based 16‐bit fixed‐point accelerator with SVD for CNN on 32‐bit memory‐limited systems

  • Feb 13, 2024
  • International Journal of Circuit Theory and Applications
  • Rama Muni Reddy Yanamala +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.