• Home
  • Search
  • FuseKNA: Fused Kernel Convolution based Accelerator for Deep Neural Networks
  • Cite Icon20
  • https://doi.org/10.1109/hpca51647.2021.00079Copy DOI Icon

FuseKNA: Fused Kernel Convolution based Accelerator for Deep Neural Networks

  • Feb 1, 2021
  • Jianxun Yang +6 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Bit-serial computation has been a prevailing convolution method to accelerate varying-precision DNNs by slicing a multi-bit data into multiple 1-bit data and transforming a multiplication into multiple additions, where additions of zero bits are ineffectual, while additions of non-zero bits are repetitive since multiple kernels are quite possible to possess non-zero bits at the same kernel positions. Previous bit-serial accelerators only remove ineffectual additions by skipping computation of zero bits, however, repetitive additions are unable to be eliminated since they compute convolution of each kernel independently. In this work, we propose fused kernel convolution algorithm to eliminate both ineffectual and repetitive additions in bit-serial computation by exploiting bit repetition and bit sparsity in weights, for both convolutional and fully-connected layers. It unifies convolutions of multiple kernels into convolution of one fused kernel by firstly grouping additions into different patterns and secondly reconstructing convolution results, minimizing addition count. Meantime, the memory accesses of activations and partial sums are decreased due to less convolution count. Then a fused kernel convolution based accelerator, FuseKNA, is designed with compact compute logic, which fully exploits value sparsity of activations and bit sparsity of weights. Benchmarked with a set of mainstream DNNs, FuseKNA improves performance by $4.47 \times$, $2.31 \times$ and $1.81 \times$, energy efficiency by $4.13 \times$, $3.06 \times$ and $2.53 \times$ over state-of-the-art Stripes, Pragmatic and Bit-Tactical.

Similar Papers
  • Conference Article
  • Citations554

Timeloop: A Systematic Approach to DNN Accelerator Evaluation

  • Mar 01, 2019
  • Angshuman Parashar +9
  • Research Article
  • Citations5

A 90.7-nW Vibration-Based Condition Monitoring Chip Featuring a Digital Compute-in-Memory- Based DNN Accelerator Using an Ultra-Low-Power 13T-SRAM Cell

  • Jan 01, 2025
  • IEEE Journal of Solid-State Circuits
  • Haochen Zhang +6
  • Conference Article
  • Citations18

Fault-free: A Fault-resilient Deep Neural Network Accelerator based on Realistic ReRAM Devices

  • Dec 05, 2021
  • Hyein Shin +2
  • Research Article
  • Citations209

MAERI

  • Mar 19, 2018
  • ACM SIGPLAN Notices
  • Hyoukjun Kwon +2
  • Research Article
  • Citations7

NeuroSpector: Systematic Optimization of Dataflow Scheduling in DNN Accelerators

  • Aug 01, 2023
  • IEEE Transactions on Parallel and Distributed Systems
  • Chanho Park +3
  • Conference Article
  • Citations5

Dynamic Mapping Mechanism to Compute DNN Models on a Resource-limited NoC Platform

  • Apr 19, 2021
  • Kun-Chih Jimmy Chen +3
  • Research Article
  • Citations18

Swallow: A Versatile Accelerator for Sparse Neural Networks

  • Mar 06, 2020
  • IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
  • Bosheng Liu +3
  • Conference Article

GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN acceleration

  • Aug 06, 2025
  • Jordi Fornt +7
  • Conference Article
  • Citations2

A Low-cost High-performance 2D-Convolution Accelerator for Deep Neural Networks in IoT

  • Dec 20, 2022
  • Hung K Nguyen +1
  • Research Article
  • Citations12

Optimizing Off-Chip Memory Access for Deep Neural Network Accelerator

  • Apr 01, 2022
  • IEEE Transactions on Circuits and Systems II: Express Briefs
  • Yong Zheng +4
  • Research Article
  • Citations22

Neural Synaptic Plasticity-Inspired Computing: A High Computing Efficient Deep Convolutional Neural Network Accelerator

  • Dec 21, 2020
  • IEEE Transactions on Circuits and Systems I: Regular Papers
  • Zihan Xia +4
  • Conference Article
  • Citations16

SCALENet

  • May 30, 2018
  • Colin Shea +2
  • Research Article
  • Citations24

Data multiplexed and hardware reused architecture for deep neural network accelerator

  • Nov 15, 2021
  • Neurocomputing
  • Gopal Raut +5
  • Conference Article
  • Citations1

Enabling Resistive-RAM-based Activation Functions for Deep Neural Network Acceleration

  • Sep 07, 2020
  • Zihan Zhang +7
  • Conference Article
  • Citations1

A High Efficiency Accelerator for Deep Neural Networks

  • Mar 01, 2018
  • Aliasger Zaidy +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.