• Open Access IconOpen Access
  • Cite Icon135
  • https://doi.org/10.1145/3289602.3293902Copy DOI Icon

Synetgy

  • Feb 20, 2019
  • Yifan Yang +10 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Using FPGAs to accelerate ConvNets has attracted significant attention in recent years. However, FPGA accelerator design has not leveraged the latest progress of ConvNets. As a result, the key application characteristics such as frames-per-second (FPS) are ignored in favor of simply counting GOPs, and results on accuracy, which is critical to application success, are often not even reported. In this work, we adopt an algorithm-hardware co-design approach to develop a ConvNet accelerator called Synetgy and a novel ConvNet model called DiracDeltaNet$^{\dagger}$. Both the accelerator and ConvNet are tailored to FPGA requirements. DiracDeltaNet, as the name suggests, is a ConvNet with only $1\times 1$ convolutions while spatial convolutions are replaced by more efficient shift operations. DiracDeltaNet achieves competitive accuracy on ImageNet (88.7\% top-5), but with 42$\times$ fewer parameters and 48$\times$ fewer OPs than VGG16. We further quantize DiracDeltaNet's weights to 4-bit and activations to 4-bits, with less than 1\% accuracy loss. These quantizations exploit well the nature of FPGA hardware. In short, DiracDeltaNet's small model size, low computational OP count, low precision and simplified operators allow us to co-design a highly customized computing unit for an FPGA. We implement the computing units for DiracDeltaNet on an Ultra96 SoC system through high-level synthesis. Our accelerator's final top-5 accuracy of 88.1\% on ImageNet, is higher than all the previously reported embedded FPGA accelerators. In addition, the accelerator reaches an inference speed of 66.3 FPS on the ImageNet classification task, surpassing prior works with similar accuracy by at least 11.6$\times$.

Similar Papers
  • Research Article
  • Citations9

Mitigating Memory-Induced Dark Silicon in Many-Accelerator Architectures

  • Jul 01, 2015
  • IEEE Computer Architecture Letters
  • Dionysios Diamantopoulos +3
  • PDF
  • Research Article
  • Citations34

Real-Time Driving Scene Semantic Segmentation

  • Jan 01, 2020
  • IEEE Access
  • Wenfu Wang +4
  • Research Article
  • Citations282

Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies

  • Aug 01, 1998
  • The VLDB Journal The International Journal on Very Large Data Bases
  • Soumen Chakrabarti +3
  • Conference Article
  • Citations4

Accelerating NLP Tasks on FPGA with Compressed BERT and a Hardware-Oriented Early Exit Method

  • Jul 01, 2022
  • Binjing Li +3
  • Conference Article
  • Citations8

Low Latency Spiking ConvNets with Restricted Output Training and False Spike Inhibition

  • Jul 01, 2018
  • Ruizhi Chen +5
  • Research Article

Railway Platform Clearance Measurement Method Based on Monocular Computer Vision Assisted with Structured Light

  • Jul 01, 2022
  • Journal of Physics: Conference Series
  • Man Liang +3
  • PDF
  • Research Article
  • Citations16

MIOpen: An Open Source Library For Deep Learning Primitives

  • Dec 17, 2020
  • Proceedings of the 30th International Conference on Computer Graphics and Machine Vision (GraphiCon 2020). Part 2
  • Jehandad Khan +14
  • Research Article
  • Citations60

Memory- and Communication-Aware Model Compression for Distributed Deep Learning Inference on IoT

  • Oct 08, 2019
  • ACM Transactions on Embedded Computing Systems
  • Kartikeya Bhardwaj +3
  • Conference Article
  • Citations29

OverGen: Improving FPGA Usability through Domain-specific Overlay Generation

  • Oct 01, 2022
  • Sihao Liu +11
  • Research Article
  • Citations5

Double-precision Dual Mode Logic carry-save multiplier

  • Aug 24, 2018
  • Integration
  • Raffaele De Rose +2
  • Research Article
  • Citations1

The Research of Gas Sealing Performance Inspection Module of Electronic Sphygmomanometer Based on STC12C5A

  • Sep 01, 2013
  • Advanced Materials Research
  • Xue Zhe Li +1
  • Research Article

HIM‐PyraNet: Hierarchical Attention and Region‐Focused Lightweight Network for Micro‐Expression Recognition

  • Dec 03, 2025
  • Concurrency and Computation: Practice and Experience
  • Fangjie Xue +4
  • Conference Article
  • Citations13

Reduced Precision Strategies for Deep Learning: A High Energy Physics Generative Adversarial Network Use Case

  • Jan 01, 2021
  • Florian Rehm +7
  • Research Article
  • Citations9

Enhanced Performance Stabilization Increases Performance Variability in a Virtual Interception Task.

  • Sep 16, 2020
  • Perceptual and Motor Skills
  • Crislaine Rangel Couto +7
  • Research Article
  • Citations122

Iron-Loss Modeling for Rotating Machines: Comparison Between Bertotti's Three-Term Expression and 3-D Eddy-Current Analysis

  • Aug 01, 2010
  • IEEE Transactions on Magnetics
  • Katsumi Yamazaki +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.