• Home
  • Search
  • Effectiveness of trace sampling for performance debugging tools
  • Open Access IconOpen Access
  • Cite Icon58
  • https://doi.org/10.1145/166955.167023Copy DOI Icon

Effectiveness of trace sampling for performance debugging tools

  • Jun 1, 1993
  • Margaret Martonosi +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Recently there has been a surge of interest in developing performance debugging tools to help programmers tune their applications for better memory performance [2, 4, 10]. These tools vary both in the detail of feedback provided to the user, and in the run-time overbead of using them. MemSpy [10] is a simulation-based tool which gives programmers detailed statistics on the memory system behavior of applications. It provides information on the frequency and causes of cache misses, and presents it in terms of source-level data and code objects with which the programmer is familiar. However, using MemSpy increases a program's execution time by roughly 10 to 40 fold. This overhead is generally acceptable for applications with execution times of several minutes or less, but it can be inconvenient when tuning applications with very long execution times.This paper examines the use of trace sampling techniques to reduce the execution time overhead of tools like MemSpy. When simulating one tenth of the references, we find that MemSpy's execution time overhead is improved by a factor of 4 to 6. That is, the execution time when using MemSpy is generally within a factor of 3 to 8 times the normal exwution time. With this improved performance, we observe only small errors in the performance statistics reported by MemSpy. On moderate sized caches of 16KB to 128KB, simulating as few as one tenth of the references (in samples of 0.5M references each) allows us to estimate the program's actual cache miss rate with an absolute error no greater than 0.3% on our five benchmarks. These errors are quite tolerable within the context of performance bugging. With larger caches we can also obtain good accuracy by using longer sample lengths. We conclude that, used with care, trace sampling is a powerful technique that makes possible performance debugging tools which provide both detailed memory statistics and low execution time overheads.

Similar Papers
  • Research Article
  • Citations1

Joint variable partitioning and bank selection instruction optimization for partitioned memory architectures

  • Mar 10, 2013
  • ACM Transactions on Embedded Computing Systems
  • Tiantian Liu +2
  • Conference Article
  • Citations28

Split-stream dictionary program compression

  • May 01, 2000
  • Steven Lucco
  • Conference Article
  • Citations2

A methodology to compute task remaining execution time

  • Jan 01, 2004
  • S Tasneem +2
  • Research Article
  • Citations39

Modeling Control Speculation for Timing Analysis

  • Jan 01, 2005
  • Real-Time Systems
  • Xianfeng Li +2
  • Research Article
  • Citations20

Resource provisioning in scalable cloud using bio-inspired artificial neural network model

  • Nov 05, 2020
  • Applied Soft Computing
  • Pradeep Singh Rawat +3
  • Conference Article
  • Citations9

State Design Pattern Implementation of a DSP processor: A case study of TMS5416C

  • Jun 01, 2011
  • Tanin Afacan
  • Conference Article
  • Citations237

The impact of architectural trends on operating system performance

  • Jan 01, 1995
  • M Rosenblum +4
  • Research Article
  • Citations18

A model of memory contention in a paging machine

  • Aug 01, 1972
  • Communications of the ACM
  • P H Oden +1
  • Book Chapter
  • Citations8

A Cache Simulator for Shared Memory Systems

  • Jan 01, 2001
  • Florian Schintke +2
  • Conference Article
  • Citations27

JVM fuzzing for JIT-induced side-channel detection

  • Jun 27, 2020
  • Tegan Brennan +2
  • Research Article
  • Citations58

On applying or-parallelism and tabling to logic programs

  • Jan 01, 2005
  • Theory and Practice of Logic Programming
  • Ricardo Rocha +2
  • Conference Article

Java bytecode optimizations

  • Feb 23, 1997
  • H.D Lambridge
  • Conference Article
  • Citations150

A technique for dynamic updating of Java software

  • Oct 03, 2002
  • A Orso +2
  • Conference Article
  • Citations1

EJOP: An Extensible Java Processor with Reasonable Performance/Flexibility Trade-off

  • Sep 01, 2012
  • Samaneh Talebi +2
  • Conference Article
  • Citations28

Decentralized Privacy-Preserving Timed Execution in Blockchain-Based Smart Contract Platforms

  • Dec 01, 2018
  • Chao Li +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.