• Home
  • Search
  • A configurable SIMD architecture with explicit datapath for intelligent learning
  • Cite Icon13
  • https://doi.org/10.1109/samos.2016.7818343Copy DOI Icon

A configurable SIMD architecture with explicit datapath for intelligent learning

  • Jul 1, 2016
  • Yifan He +7 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The use of a wide Single-Instruction-Multiple-Data (SIMD) architecture is a promising approach to build energy-efficient high performance embedded processors. In this paper, based on our design framework for low-power SIMD processors, we propose a multiply-accumulate (MAC) unit with variable number of accumulator registers. The proposed MAC unit exploits both the merits of merged operation and register tiling. A Convolutional Neural Network (CNN) is a popular learning based algorithm due to its flexibility and high accuracy. However, a CNN-based application is often computationally intensive as it applies convolution operations extensively on a large data set. In this work, a CNN-based intelligent learning application is analyzed and mapped in the context of SIMD architectures. Experimental results show that the proposed architecture is efficient. In a 64-PE instance, the proposed SIMD processor with MAC4reg achieves an effective performance of 63.2 GOPS. Compared to the two baseline SIMD processors without MAC4reg, the proposed design brings 54.0% and 32.1% reduction in execution time, and 20.5% and 35.1% reduction in energy consumption respectively.

Similar Papers
  • Book Chapter

Performance Analysis of Existing SIMD Architectures

  • Jan 01, 2019
  • Chao Cui +2
  • Conference Article

Accelearation of Full-Search Algorithm on SIMD Architectures by Using Eight-Bit Partial Sums of Four Luminance Values

  • Dec 01, 2006
  • C J Duanmu
  • Conference Article
  • Citations3

On the scalability of SIMD processing for software defined radio algorithms

  • Jul 01, 2010
  • Peter Westermann +1
  • Research Article
  • Citations2

A MATLAB Vectorizing Compiler Targeting Application-Specific Instruction Set Processors

  • Jan 04, 2017
  • ACM Transactions on Design Automation of Electronic Systems
  • Ioannis Latifis +6
  • Research Article
  • Citations11

A Low-Energy Wide SIMD Architecture with Explicit Datapath

  • Sep 16, 2014
  • Journal of Signal Processing Systems
  • Luc Waeijen +3
  • Research Article
  • Citations37

Low Complexity Multiply Accumulate Unit for Weight-Sharing Convolutional Neural Networks

  • Jul 01, 2017
  • IEEE Computer Architecture Letters
  • James Garland +1
  • Conference Article

Optimization of H.264 video decoder from Android OS for MIPS32 DSP ASE architectures

  • Nov 01, 2015
  • Branimir Vasić +3
  • Conference Article
  • Citations5

MAC unit for reconfigurable systems using multi-operand adders with double carry-save encoding

  • Apr 01, 2016
  • Ugur Cini +1
  • Conference Article
  • Citations21

A Systolic Dataflow Based Accelerator for CNNs

  • Oct 01, 2020
  • Saptarsi Das +4
  • Research Article
  • Citations5

Real-time Face Recognition Using SIMD and VLIW Architecture

  • Jan 01, 2006
  • Journal of Computing and Information Technology
  • Srinivasa Kumar Devireddy +2
  • Conference Article
  • Citations7

Design & implementation of area efficient low power high speed MAC unit using FPGA

  • Sep 01, 2017
  • Roshani Pawar +1
  • Research Article
  • Citations67

A High-Speed, Energy-Efficient Two-Cycle Multiply-Accumulate (MAC) Architecture and Its Application to a Double-Throughput MAC Unit

  • Dec 01, 2010
  • IEEE Transactions on Circuits and Systems I: Regular Papers
  • Tung Thanh Hoang +2
  • Research Article
  • Citations2

A High-Throughput Multiply-Accumulate Unit With Long Feedback Loop Using Low-Voltage Rapid Single-Flux Quantum Circuits

  • Apr 01, 2023
  • IEEE Transactions on Applied Superconductivity
  • Ikki Nagaoka +7
  • Conference Article
  • Citations16

SCALENet

  • May 30, 2018
  • Colin Shea +2
  • PDF
  • Research Article
  • Citations7

Arithmetic Coding-Based 5-Bit Weight Encoding and Hardware Decoder for CNN Inference in Edge Devices

  • Jan 01, 2021
  • IEEE Access
  • Jong Hun Lee +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.