• Home
  • Search
  • Deadline-Aware MapReduce Job Scheduling with Dynamic Resource Availability
  • Cite Icon34
  • https://doi.org/10.1109/tpds.2018.2873373Copy DOI Icon

Deadline-Aware MapReduce Job Scheduling with Dynamic Resource Availability

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

As MapReduce is becoming ubiquitous in large-scale data analysis, many recent studies have shown that the performance of MapReduce could be improved by different job scheduling approaches, e.g., Fair Scheduler and Capacity Scheduler. However, most exiting MapReduce job schedulers focus on the scenario that MapReduce cluster is stable and pay little attention to the MapReduce cluster with dynamic resource availability. In fact, MapReduce cluster resources may fluctuate as there is a growing number of Hadoop clusters deployed on hybrid systems, e.g., infrastructure powered by mix of traditional and renewable energy, and cloud platforms hosting heterogeneous workloads. Thus, there is a growing need for providing predictable services to users who have strict requirements on job completion times in such dynamic environments. In this paper, we propose, RDS , a Resource and Deadline-aware Hadoop job Scheduler that takes future resource availability into consideration when minimizing job deadline misses. We formulate the job scheduling problem as an online optimization problem and solve it using an efficient receding horizon control algorithm. To aid the control, we design a self-learning model to estimate job completion times. We further extend the design of RDS scheduler to support flexible performance goals in various dynamic clusters. In particular, we use flexible deadline time bounds instead of the single fixed job completion deadline. We have implemented RDS in the open-source Hadoop implementation and performed evaluations with various benchmark workloads. Experimental results show that RDS substantially reduces the penalty of deadline misses by at least 36 and 10 percent compared with Fair Scheduler and Earliest Deadline First (EDF) scheduler, respectively. In a Hadoop cluster running partially on renewable energy, the experimental result shows the green power based resource prediction approach can further reduce the penalty of deadline misses by 16 percent compared to Auto-Regressive Integrated Moving Average (ARIMA) prediction approach.

Similar Papers
  • Conference Article
  • Citations22

Task-Cloning Algorithms in a MapReduce Cluster with Competitive Performance Bounds

  • Jun 01, 2015
  • Huanle Xu +1
  • Research Article

Application and Storage-Aware Data Placement and Job Scheduling for Hadoop Clusters

  • Dec 21, 2020
  • Journal of Circuits, Systems and Computers
  • Tao Li +5
  • Book Chapter
  • Citations5

Adaptive Job Scheduling for a Service Grid Using a Genetic Algorithm

  • Jan 01, 2004
  • Yang Gao +4
  • Research Article
  • Citations1

Online Job Dispatching and Scheduling to Minimize Job Completion Time and to Meet Deadlines

  • Dec 01, 2018
  • Journal of Interconnection Networks
  • Yupeng Li
  • Book Chapter

Enhancing the Performance of MapReduce Default Scheduler by Detecting Prolonged TaskTrackers in Heterogeneous Environments

  • Sep 04, 2015
  • Nenavath Srinivas Naik +2
  • Conference Article

On-line batch scheduling on tow parallel machines with wait

  • Dec 01, 2015
  • Huo Manchen +1
  • Conference Article
  • Citations11

Performance evaluation of fair and capacity scheduling in Hadoop YARN

  • Oct 01, 2015
  • Garima Sharma +1
  • Research Article
  • Citations30

Job scheduling in mesh multicomputers

  • Jan 01, 1998
  • IEEE Transactions on Parallel and Distributed Systems
  • D Das Sharma +1
  • Conference Article
  • Citations1

Fair and Delay Adaptive Scheduler (FDAS) preliminary modeling and optimization

  • Jan 01, 2016
  • Abdelwahab M Elnaka +2
  • Research Article
  • Citations20

Optimizing makespan and resource utilization for multi-DNN training in GPU cluster

  • Jun 24, 2021
  • Future Generation Computer Systems
  • Zhongjin Li +5
  • Conference Article
  • Citations1

Batch computer scheduling

  • Jan 01, 1974
  • Stephen R Kimbleton
  • Research Article
  • Citations50

Scheduling Jobs across Geo-Distributed Datacenters with Max-Min Fairness

  • Jul 01, 2019
  • IEEE Transactions on Network Science and Engineering
  • Li Chen +3
  • Conference Article
  • Citations115

Evaluating the impact of job scheduling and power management on processor lifetime for chip multiprocessors

  • Jun 15, 2009
  • Ayse K Coskun +3
  • Conference Article
  • Citations76

Cost-Wait Trade-Offs in Client-Side Resource Provisioning with Elastic Clouds

  • Apr 26, 2011
  • Stephane Genaud +1
  • Research Article
  • Citations8

P4INC-AOI: All-Optical Interconnect Empowered by In-Network Computing for DML Workloads

  • Jun 01, 2025
  • IEEE Transactions on Networking
  • Xuexia Xie +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.