• Home
  • Search
  • Decoupling local variable accesses in a wide-issue superscalar processor
  • Cite Icon29
  • https://doi.org/10.1145/307338.300988Copy DOI Icon

Decoupling local variable accesses in a wide-issue superscalar processor

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Providing adequate data bandwidth is extremely important for a wide-issue superscalar processor to achieve its full performance potential. Adding a large number of ports to a data cache, however, becomes increasingly inefficient and can add to the hardware complexity significantly. This paper takes an alternative or complementary approach for providing more data bandwidth, called the data-decoupled architecture. The approach, with support from the compiler and/or hardware, partitions the memory stream into two independent streams early in the processor pipeline, and feeds each stream to a separate memory access queue and cache. Under this model, the paper studies the potential of decoupling memory accesses to program's local variables that are allocated on the run-time stack. Using a set of integer and floating-point programs from the SPEC95 benchmark suite, it is shown that local variable accesses constitute a large portion of all the memory references, while their reference space is very small, averaging around 7 words per (static) procedure. To service local variable accesses quickly, two optimizations, fast data forwarding and access combining, are proposed and studied. Some of the important design parameters, such as the cache size, the number of cache ports, and the degree of access combining, are studied based on simulations. The potential performance of the proposed scheme is measured using various configurations, and it is concluded that the scheme can become a viable alternative to building a single multi-ported data cache.

Similar Papers
  • Research Article

THE DESIGN OF THE PIPELINED RISC-V PROCESSOR WITH THE HARDWARE COPROCESSOR OF DIGITAL SIGNAL PROCESSING

  • Apr 02, 2024
  • Radio Electronics Computer Science Control
  • Y Y Vavruk +2
  • Research Article
  • Citations3

Architectural Support for Coherent Architecturally Visible Storage in Instruction Set Extensions

  • Jan 01, 2010
  • Infoscience (Ecole Polytechnique Fédérale de Lausanne)
  • Ties Jan Henderikus Kluter
  • Research Article
  • Citations2

Estimation of Web Proxy Server Cache Size using G/G/1 Queuing Model

  • Jan 01, 2010
  • International Journal of Applied Research on Information Technology and Computing
  • Riktesh Srivastava
  • Book Chapter
  • Citations1

Dynamic Partition of Memory Reference Instructions – A Register Guided Approach

  • Jan 01, 2005
  • Yixin Shi +1
  • Conference Article
  • Citations10

Custom-sized caches in application-specific memory hierarchies

  • Dec 01, 2015
  • Felix Winterstein +4
  • Conference Article
  • Citations33

A fast analytical model of fully associative caches

  • Jun 08, 2019
  • Tobias Gysi +3
  • Conference Article

Efficient address generation for affine subscripts in data-parallel programs

  • Dec 14, 1998
  • Kuei-Ping Shih +2
  • Conference Article
  • Citations16

Hybrid source-level simulation of data caches using abstract cache models

  • Mar 01, 2012
  • S Stattelmann +4
  • Conference Article
  • Citations2

Improving reliable data transport in wireless sensor networks through dynamic cache-aware rate control mechanism

  • Oct 01, 2017
  • Melchizedek I Alipio +1
  • Book Chapter
  • Citations3

A Cache-Aware Congestion Control for Reliable Transport in Wireless Sensor Networks

  • Jan 01, 2018
  • Lecture notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering
  • Melchizedek I Alipio +1
  • Research Article

Enabling MPEG-2 video playback in embedded systems through improved data cache efficiency

  • Feb 01, 2006
  • IEEE Transactions on Multimedia
  • P Soderquist +2
  • Conference Article
  • Citations64

Impact of Cache Architecture and Interface on Performance and Area of FPGA-Based Processor/Parallel-Accelerator Systems

  • Apr 01, 2012
  • Jongsok Choi +5
  • Research Article
  • Citations185

Optimizing main-memory join on modern hardware

  • Jan 01, 1999
  • IEEE Transactions on Knowledge and Data Engineering
  • S Manegold +2
  • Conference Article
  • Citations19

Memory access patterns of occlusion-compatible 3D image warping

  • Jan 01, 1997
  • William R Mark +1
  • Conference Article
  • Citations2

Introducing Energy Efficiency into Graphics Processors

  • Dec 01, 2010
  • B V N Silpa +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.