• Home
  • Search
  • Instruction scheduling for instruction level parallel processors
  • Open Access IconOpen Access
  • Cite Icon58
  • https://doi.org/10.1109/5.964443Copy DOI Icon

Instruction scheduling for instruction level parallel processors

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Nearly all personal computer and workstation processors, and virtually all high-performance embedded processor cores, now embody instruction level parallel (ILP) processing in the form of superscalar or very long instruction word (VLIW) architectures. ILP processors put much more of a burden on compilers; without heroic compiling techniques, most such processors fall far short of their performance goals. Those techniques are largely found in the high-level optimization phase and in the code generation phase; they are also collectively called instruction scheduling. This paper reviews the state of the art in code generation for ILP parallel processors. Modern ILP code generation methods move code across basic block boundaries. These methods grew out of techniques for generating horizontal microcode, so we introduce the problem by describing its history. Most modem approaches can be categorized by the shape of the scheduling region. Some of these regions are loops, and for those techniques known broadly as are used. Software Pipelining techniques are only considered here when there are issues relevant to the region-based techniques presented. The selection of a type of region to use in this process is one of the most controversial questions in code generation; the paper surveys the best known alternatives. The paper then considers two questions: First, given a type of region, how does one pick specific regions of that type in the intermediate code. In conjunction with region selection, we consider region enlargement techniques such as unrolling and branch target expansion. The second question, how does one construct a schedule once regions have been selected, occupies the next section of the paper. Finally, schedule construction using recent, innovative resource modeling based on finite-state automata is then reexamined. The paper includes an extensive bibliography.

Similar Papers
  • Conference Article
  • Citations9

Tetris

  • Jun 13, 2007
  • Weifeng Xu +1
  • Conference Article
  • Citations3

Exploiting Java instruction/thread level parallelism with horizontal multithreading

  • Jan 15, 2001
  • Kenji Watanabe +2
  • Conference Article
  • Citations4

Performance analysis of inter cluster communication methods in VLIW architecture

  • Jan 05, 2004
  • S Saluja +1
  • Conference Article

3D graphics system with VLIW processor for geometry acceleration

  • Aug 28, 2000
  • Young-Wook Jeon +6
  • Book Chapter
  • Citations18

Automatically Customising VLIW Architectures with Coarse Grained Application-Specific Functional Units

  • Jan 01, 2004
  • Diviya Jain +3
  • Research Article
  • Citations30

Exploiting instruction-level parallelism for integrated control-flow monitoring

  • Jan 01, 1994
  • IEEE Transactions on Computers
  • M.A Schuette +1
  • Conference Article
  • Citations2

Improving performance in VLIW soft-core processors through software-controlled scratchpads

  • Jul 01, 2016
  • Tiago Jost +2
  • Conference Article
  • Citations1

A Holistic Approach to CPU Verification using Formal Techniques

  • Jun 02, 2022
  • Dasari Bhavya Sri +3
  • Book Chapter
  • Citations2

Rule-Based Power-Balanced VLIW Instruction Scheduling with Uncertainty

  • Jan 01, 2005
  • Shu Xiao +2
  • Conference Article

Global instruction scheduling in dynamic compilation for embedded systems

  • Jan 01, 2006
  • Giovanni Agosta +3
  • Dissertation
  • Citations3

Loop transformations for clustered VLIW architectures

  • Jan 01, 2002
  • Yi Qian
  • Research Article
  • Citations1

Single instruction multiple data code auto generation for a very long instruction words digital signal processor in sensor‐based systems

  • Jun 01, 2013
  • IET Wireless Sensor Systems
  • Xu Yang +4
  • Dissertation

Improving multithreading performance for clustered VLIW architectures.

  • Jun 14, 2013
  • Manoj Gupta
  • Book Chapter
  • Citations15

Code Generation for STA Architecture

  • Jan 01, 2006
  • J Guo +5
  • Research Article

INCORPORATING FAULT-TOLERANT FEATURES IN VLIW PROCESSORS

  • Oct 01, 2005
  • International Journal of Reliability, Quality and Safety Engineering
  • Yung-Yuan Chen
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.