• Home
  • Search
  • A Graph-Based Learning Framework for Compiler Loop Auto-Vectorization
  • https://doi.org/10.34133/icomputing.0113Copy DOI Icon

A Graph-Based Learning Framework for Compiler Loop Auto-Vectorization

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

The single instruction multiple data (SIMD) capability in modern processors is critical to improving the performance of current compute-intensive programs. Modern compilers use vectorization techniques to exploit the SIMD capability, by detecting data parallelism in scalar source code and transforming a group of scalar instructions into vector-based instructions. In this study, we focus on one of the most common vectorization techniques, a technique called loop-based vectorization, which targets loops and optimizes their performance by grouping multiple occurrences of the same operation across loop iterations into a single SIMD instruction. We propose a data-driven graph-based learning framework for automatic vectorization, called autograph , which takes an input program, extracts the loops, and then learns a structured representation to automatically predict the correct vectorization and interleaving factors. Our proposed framework utilizes deep reinforcement learning to learn an optimal policy (observations to actions) from an intelligent agent in a SIMD environment, and automatically injects the predicted vectorization pragmas into the input program. We conducted an extensive evaluation on multiple benchmark datasets and comparisons with state-of-the-art baselines. Our results show that autograph achieves on average 2.49× performance improvement for Polybench compared to NeuroVectorizer and 3.69× compared to the baseline -O3.

Similar Papers
  • Conference Article
  • Citations6

Customized SIMD unit synthesis for system on programmable chip

  • Jan 01, 2006
  • Muhammad Omer Cheema +1
  • Conference Article
  • Citations3

Reduction of Complexity and Automation of Parallel Execution through Loop Level Parallelism

  • Jan 01, 2007
  • Robert A Tefft +1
  • Research Article
  • Citations1

Single instruction multiple data code auto generation for a very long instruction words digital signal processor in sensor‐based systems

  • Jun 01, 2013
  • IET Wireless Sensor Systems
  • Xu Yang +4
  • Conference Article
  • Citations104

From SODA to scotch: The evolution of a wireless baseband processor

  • Nov 01, 2008
  • Mark Woh +10
  • Book Chapter

Performance Analysis of Existing SIMD Architectures

  • Jan 01, 2019
  • Chao Cui +2
  • Research Article
  • Citations2

Exploiting SIMD-Ified Bit-Parallelism for High-Performance Complex Event Matching

  • Feb 01, 2026
  • IEEE Transactions on Knowledge and Data Engineering
  • Tao Qiu +5
  • PDF
  • Conference Article
  • Citations41

VeGen: a vectorizer generator for SIMD and beyond

  • Apr 17, 2021
  • Yishen Chen +3
  • Conference Article
  • Citations39

Introducing Control Flow into Vectorized Code

  • Sep 15, 2007
  • Jaewook Shin
  • Research Article

Automation of program code vectorization for the use of SIMD instructions

  • Jan 01, 2026
  • Mathematical machines and systems
  • M.O Bohdiuk +1
  • Conference Article
  • Citations3

On the scalability of SIMD processing for software defined radio algorithms

  • Jul 01, 2010
  • Peter Westermann +1
  • Conference Article

Accelearation of Full-Search Algorithm on SIMD Architectures by Using Eight-Bit Partial Sums of Four Luminance Values

  • Dec 01, 2006
  • C J Duanmu
  • Conference Article
  • Citations2

Accelerating higher-order masking of AES using composite field and SIMD

  • Dec 01, 2015
  • Abdulaziz Miyajan +3
  • Research Article
  • Citations8

Performance of an advanced video codec on a general-purpose processor with media ISA extensions

  • Jan 01, 2000
  • IEEE Transactions on Consumer Electronics
  • V Lappalainen
  • Conference Article
  • Citations5

Design of a custom vector operation API exploiting SIMD intrinsics within Java

  • May 01, 2010
  • Jonathan Parri +4
  • Conference Article
  • Citations1

Optimizing Bandwidth Constraint through Register Interconnection for Stream Processors

  • Sep 15, 2007
  • Weihua Zhang +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.