• Home
  • Search
  • Coded Computation Over Heterogeneous Clusters
  • Open Access IconOpen Access
  • Cite Icon238
  • https://doi.org/10.1109/tit.2019.2904055Copy DOI Icon

Coded Computation Over Heterogeneous Clusters

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In large-scale distributed computing clusters, such as Amazon EC2, there are several types of “system noise” that can result in major degradation of performance: system failures, bottlenecks due to limited communication bandwidth, latency due to straggler nodes, and so on. There have been recent results that demonstrate the impact of coding for efficient utilization of computation and storage redundancy to alleviate the effect of stragglers and communication bottlenecks in homogeneous clusters. In this paper, we focus on general heterogeneous distributed computing clusters consist of a variety of computing machines with different capabilities. We propose a coding framework for speeding up distributed computing in heterogeneous clusters by trading redundancy for reducing the latency of computation. In particular, we propose heterogeneous coded matrix multiplication (HCMM) algorithm for performing distributed matrix multiplication over heterogeneous clusters that are provably asymptotically optimal for a broad class of processing time distributions. Moreover, we show that HCMM is unboundedly faster than any uncoded scheme that partitions the total workload among the workers. To demonstrate how the proposed HCMM scheme can be applied in practice, we provide results from numerical studies and Amazon EC2 experiments comparing HCMM with three benchmark load allocation schemes—uniform uncoded, load-balanced uncoded, and uniform coded. In particular, in our numerical studies, HCMM achieves speedups of up to 73%, 56%, and 42%, respectively, over the three benchmark schemes mentioned earlier. Furthermore, we carry out experiments over Amazon EC2 clusters and demonstrate how HCMM can be combined with rateless codes with nearly linear decoding complexity. In particular, we show that HCMM combined with the Luby transform codes can significantly reduce the overall execution time. HCMM is found to be up to 61%, 46%, and 36% faster than the aforementioned three benchmark schemes, respectively. Additionally, we provide a generalization to the problem of optimal load allocation in heterogeneous settings, where we take into account the monetary costs associated with distributed computing clusters. We argue that HCMM is asymptotically optimal for budget-constrained scenarios as well. In particular, we characterize the minimum possible expected cost associated with a computation task over a given cluster of machines. Furthermore, we develop a heuristic algorithm for (HCMM) load allocation for the distributed implementation of budget-limited computation tasks.

Similar Papers
  • Conference Article
  • Citations17

Energy-Efficient Delay Time-Based Process Allocation Algorithm for Heterogeneous Server Clusters

  • Mar 01, 2015
  • Tomoya Enokido +1
  • Conference Article

Efficient broadcasts and simple algorithms for parallel linear algebra computing in clusters

  • Apr 22, 2003
  • F.G Tinetti +1
  • Conference Article
  • Citations1

The Improved Redundant Delay Time-Based (IRDTB) Algorithm to Perform Computation Type Application Processes in Heterogeneous Server Clusters

  • Sep 01, 2016
  • Tomoya Enokido +1
  • Research Article
  • Citations17

Programming Heterogeneous Clusters with Accelerators Using Object-Based Programming

  • Jan 01, 2011
  • Scientific Programming
  • David M Kunzman +1
  • Research Article
  • Citations27

TSC-based Method to Enhance Asset Utilization of Interconnected Distribution Systems

  • Jan 01, 2016
  • IEEE Transactions on Smart Grid
  • Jun Xiao +4
  • Research Article
  • Citations9

Composition of Ar–Kr, Kr–Xe, and N2–Ar Clusters Produced by Supersonic Expansion of Gas Mixtures

  • Jul 27, 2014
  • Journal of Cluster Science
  • O P Konotop +3
  • Research Article
  • Citations14

VCluster: a thread‐based Java middleware for SMP and heterogeneous clusters with thread migration support

  • Nov 21, 2007
  • Software: Practice and Experience
  • Hua Zhang +2
  • Research Article
  • Citations13

Load allocation problem for optimal design of aircraft electrical power system

  • Feb 01, 2013
  • International Journal of Applied Electromagnetics and Mechanics
  • X Giraud +6
  • Conference Article
  • Citations6

Algorithms for Optimising Heterogeneous Cloud Virtual Machine Clusters

  • Dec 01, 2016
  • Long Thai +2
  • PDF
  • Research Article
  • Citations12

Serine proteases profiles of Leishmania (Viannia) braziliensis clinical isolates with distinct susceptibilities to antimony

  • Jul 09, 2021
  • Scientific Reports
  • Anabel Zabala-Peñafiel +10
  • Research Article

Deep reinforcement learning for job scheduling on load-aware heterogeneous cluster

  • Dec 23, 2025
  • Journal of Big Data
  • Zhenjie Yao +5
  • Research Article
  • Citations3

Formation and reactions of cluster ions from aromatic carboxylic acids together with amino acids

  • Nov 01, 2001
  • Israel Journal of Chemistry
  • Anja Meffert +1
  • Conference Article
  • Citations8

Power-performance modeling of heterogeneous cluster-based web servers

  • Oct 01, 2009
  • Hiroshi Sasaki +3
  • Research Article
  • Citations28

Dynamic load balancing for I/O-intensive applications on clusters

  • Nov 01, 2009
  • ACM Transactions on Storage
  • Xiao Qin +4
  • Book Chapter
  • Citations3

Endowing the MIA Cloud Autoscaler with Adaptive Evolutionary and Particle Swarm Multi-Objective Optimization Algorithms

  • Jan 01, 2021
  • Virginia Yannibelli +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.