• Home
  • Search
  • Locality-aware task scheduling for homogeneous parallel computing systems
  • Cite Icon12
  • https://doi.org/10.1007/s00607-017-0581-6Copy DOI Icon

Locality-aware task scheduling for homogeneous parallel computing systems

  • Nov 1, 2017
  • Computing
  • Muhammad Khurram Bhatti +6 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In systems with complex many-core cache hierarchy, exploiting data locality can significantly reduce execution time and energy consumption of parallel applications. Locality can be exploited at various hardware and software layers. For instance, by implementing private and shared caches in a multi-level fashion, recent hardware designs are already optimised for locality. However, this would all be useless if the software scheduling does not cast the execution in a manner that promotes locality available in the programs themselves. Since programs for parallel systems consist of tasks executed simultaneously, task scheduling becomes crucial for the performance in multi-level cache architectures. This paper presents a heuristic algorithm for homogeneous multi-core systems called locality-aware task scheduling (LeTS). The LeTS heuristic is a work-conserving algorithm that takes into account both locality and load balancing in order to reduce the execution time of target applications. The working principle of LeTS is based on two distinctive phases, namely; working task group formation phase (WTG-FP) and working task group ordering phase (WTG-OP). The WTG-FP forms groups of tasks in order to capture data reuse across tasks while the WTG-OP determines an optimal order of execution for task groups that minimizes the reuse distance of shared data between tasks. We have performed experiments using randomly generated task graphs by varying three major performance parameters, namely: (1) communication to computation ratio (CCR) between 0.1 and 1.0, (2) application size, i.e., task graphs comprising of 50-, 100-, and 300-tasks per graph, and (3) number of cores with 2-, 4-, 8-, and 16-cores execution scenarios. We have also performed experiments using selected real-world applications. The LeTS heuristic reduces overall execution time of applications by exploiting inter-task data locality. Results show that LeTS outperforms state-of-the-art algorithms in amortizing inter-task communication cost.

Similar Papers
  • PDF
  • Research Article

Clustering-based Graph Numbering using Execution Traces for Cache Misses Reduction in Graph Analysis Applications

  • Mar 08, 2025
  • Revue Africaine de Recherche en Informatique et Mathématiques Appliquées
  • Régis Audran Mogo Wafo +3
  • Research Article
  • Citations5

PERIDOT: Modeling Execution Time of Spark Applications

  • Jan 01, 2021
  • IEEE Open Journal of the Computer Society
  • Sarah Shah +3
  • Book Chapter
  • Citations3

A Vector-Scheduling Approach for Running Many-Task Applications in the Cloud

  • Jan 01, 2018
  • Brian Peterson +3
  • Research Article
  • Citations13

Low-power high-performance reconfigurable computing cache architectures

  • Oct 01, 2004
  • IEEE Transactions on Computers
  • Rama Sangireddy +2
  • Research Article
  • Citations5

Aspect-oriented middleware framework for resolving service discovery issues in Internet of Things

  • Jan 01, 2016
  • International Journal of Internet Protocol Technology
  • Senthil Murugan Balakrishnan +1
  • Conference Article
  • Citations2

Task graph mapping and scheduling on heterogeneous architectures under communication constraints

  • Jul 01, 2017
  • Control theory & applications
  • A Emeretlis +4
  • Research Article
  • Citations34

A Bee Colony-Based Algorithm for Task Offloading in Vehicular Edge Computing

  • Sep 01, 2023
  • IEEE Systems Journal
  • Alisson Barbosa De Souza +5
  • Research Article
  • Citations3

Assessing the Suitability of King Topologies for Interconnection Networks

  • Mar 01, 2016
  • IEEE Transactions on Parallel and Distributed Systems
  • E Stafford +6
  • Conference Article
  • Citations130

Exploring the Energy-Time Tradeoff in MPI Programs on a Power-Scalable Cluster

  • Jul 15, 2015
  • V.W Freeh +4
  • Research Article
  • Citations3

Toward an SGX-Friendly Java Runtime

  • Jan 01, 2024
  • IEEE Transactions on Computers
  • Mingyu Wu +7
  • Conference Article

An evaluation of PL/I for scientific applications

  • Jan 01, 1967
  • R D Mcculloch +1
  • Research Article
  • Citations2

On the use of HCM to develop a resource allocation algorithm for heterogeneous clusters

  • Jul 31, 2018
  • Concurrency and Computation: Practice and Experience
  • Thiago Marques Soares +2
  • Research Article
  • Citations21

Locality-Aware Task Scheduling and Data Distribution for OpenMP Programs on NUMA Systems and Manycore Processors

  • Jan 01, 2015
  • Scientific Programming
  • Ananya Muddukrishna +2
  • Conference Article
  • Citations38

Controlling garbage collection and heap growth to reduce the execution time of Java applications

  • Oct 01, 2001
  • Tim Brecht +3
  • Research Article
  • Citations28

Controlling garbage collection and heap growth to reduce the execution time of Java applications

  • Sep 01, 2006
  • ACM Transactions on Programming Languages and Systems
  • Tim Brecht +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.