• Cite Icon47
  • https://doi.org/10.14778/2536222.2536225Copy DOI Icon

Piranha

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Cluster computing has emerged as a key parallel processing platform for large scale data. All major internet companies use it as their major central processing platform. One of cluster computing's most popular examples is MapReduce and its open source implementation Hadoop. These systems were originally designed for batch and massive-scale computations. Interestingly, over time their production workloads have evolved into a mix of a small fraction of large and long-running jobs and a much bigger fraction of short jobs. This came about because these systems end up being used as data warehouses, which store most of the data sets and attract ad hoc, short, data-mining queries. Moreover, the availability of higher level query languages that operate on top of these cluster systems proliferated these ad hoc queries. Since existing systems were not designed for short, latency-sensistive jobs, short interactive jobs suffer from poor response times. In this paper, we present Piranha--a system for optimizing short jobs on Hadoop without affecting the larger jobs. It runs on existing unmodified Hadoop clusters facilitating its adoption. Piranha exploits characteristics of short jobs learned from production workloads at Yahoo! clusters to reduce the latency of such jobs. To demonstrate Piranha's effectiveness, we evaluated its performance using three realistic short queries. Piranha was able to reduce the queries' response times by up to 71%.

Similar Papers
  • Conference Article
  • Citations9

Improving Short Job Latency Performance in Hybrid Job Schedulers with Dice

  • Aug 05, 2019
  • Wei Zhou +2
  • Conference Article
  • Citations11

Tyrex: Size-Based Resource Allocation in MapReduce Frameworks

  • May 01, 2016
  • Bogdan Ghit +1
  • Conference Article
  • Citations37

A Hybrid Scheduling Algorithm for Data Intensive Workloads in a MapReduce Environment

  • Nov 01, 2012
  • Phuong Nguyen +4
  • PDF
  • Research Article

Developing a Platform Using Petri Nets and GPenSIM for Simulation of Multiprocessor Scheduling Algorithms

  • Jun 29, 2024
  • Applied Sciences
  • Daniel Osmundsen Dirdal +3
  • Conference Article
  • Citations2

Comparing different scheduling schemes for M/G/1 queue

  • Dec 01, 2010
  • Sarah Tasneem +3
  • Book Chapter
  • Citations8

Load Balancing with the Help of Round Robin and Shortest Job First Scheduling Algorithm in Cloud Computing

  • Jan 01, 2021
  • Shubham Kumar +1
  • Conference Article
  • Citations1

Subject Oriented Data Partitioning – A Proposed Data Warehousing Schema

  • May 01, 2019
  • Abdul Moktadir +1
  • PDF
  • Research Article

Improved View Selection Algorithm Using SOM and 0/1 Knapsack

  • May 19, 2019
  • Statistics, Optimization & Information Computing
  • Reyhaneh Sabbagh Gol +1
  • Conference Article
  • Citations8

Cloud Native Data Platform for Network Telemetry and Analytics

  • Oct 25, 2021
  • Daniel Tovarnak +2
  • PDF
  • Research Article
  • Citations2

Performance Analysis of OS Scheduling for a Reconfigurable Computing Environment

  • Sep 16, 2015
  • Indian Journal of Science and Technology
  • B Abirami +2
  • Research Article
  • Citations100

Schema versioning in data warehouses: Enabling cross-version querying via schema augmentation

  • Oct 18, 2005
  • Data & Knowledge Engineering
  • Matteo Golfarelli +3
  • Book Chapter
  • Citations13

Chapter 9 - A Data Warehouse Strategy for on-Demand Multiscale Mapping

  • Jan 01, 2007
  • Generalisation of Geographic Information
  • Eveline Bernier +1
  • Research Article
  • Citations39

Materialised view selection using differential evolution

  • Jan 01, 2014
  • International Journal of Innovative Computing and Applications
  • T.V Vijay Kumar +1
  • PDF
  • Research Article

The realization and application of the data analysis platform of netizen behavior based on Hive

  • Oct 23, 2023
  • Applied and Computational Engineering
  • Zihao Zhao
  • Conference Article
  • Citations4

An Efficient Theta-Join Query Processing Algorithm on MapReduce Framework

  • Jun 01, 2012
  • Shih-Ying Chen +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.