• Home
  • Search
  • Improving spark application throughput via memory aware task co-location
  • Cite Icon56
  • https://doi.org/10.1145/3135974.3135984Copy DOI Icon

Improving spark application throughput via memory aware task co-location

  • Dec 11, 2017
  • Vicent Sanz Marco +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Data analytic applications built upon big data processing frameworks such as Apache Spark are an important class of applications. Many of these applications are not latency-sensitive and thus can run as batch jobs in data centers. By running multiple applications on a computing host, task co-location can significantly improve the server utilization and system throughput. However, effective task co-location is a non-trivial task, as it requires an understanding of the computing resource requirement of the co-running applications, in order to determine what tasks, and how many of them, can be co-located. State-of-the-art co-location schemes either require the user to supply the resource demands which are often far beyond what is needed; or use a one-size-fits-all function to estimate the requirement, which, unfortunately, is unlikely to capture the diverse behaviors of applications.

Similar Papers
  • PDF
  • Research Article
  • Citations9

Distributed fuzzy clustering algorithm for mixed-mode data in Apache SPARK

  • Dec 21, 2022
  • Journal of Big Data
  • Abdul Wahab Akram +1
  • Research Article
  • Citations4

The Impact of Big Data Processing Framework for Artificial Intelligence within Corporate Marketing Communication

  • Dec 13, 2018
  • International Journal of Engineering & Technology
  • Muhamad Fazil Ahmad
  • Research Article

Optimize Large-Scale Data Processing Via Spark Tuning

  • Oct 27, 2023
  • International Research Journal of Modernization in Engineering Technology and Science
  • Apurva Kumar +1
  • Research Article
  • Citations4

SEED: solar energy‐aware efficient scheduling for data centers

  • Dec 05, 2013
  • Concurrency and Computation: Practice and Experience
  • Chao Jing +2
  • Conference Article

From Network Perspective to Optimize Shuffling Phase in MapReduce

  • Aug 01, 2018
  • Haifeng Wang +1
  • Conference Article

Performance Analysis of Java Virtual Machine for Machine Learning Workloads using Apache Spark

  • Aug 25, 2016
  • N Hema +6
  • Book Chapter

Modeling Instability for Large Scale Processing Tasks Within HEP Distributed Computing Environments

  • Nov 03, 2017
  • Olga Datskova +1
  • Research Article
  • Citations83

Characteristics of Co-Allocated Online Services and Batch Jobs in Internet Data Centers: A Case Study From Alibaba Cloud

  • Jan 01, 2019
  • IEEE Access
  • Congfeng Jiang +5
  • Conference Article
  • Citations2

A Complex Task Scheduling Scheme for Big Data Platforms Based on Boolean Satisfiability Problem

  • Jul 01, 2018
  • Huang Hong +4
  • Research Article
  • Citations24

VLocality: Revisiting Data Locality for MapReduce in Virtualized Clouds

  • Jan 01, 2017
  • IEEE Network
  • Xiaoqiang Ma +4
  • Research Article
  • Citations7

A configurable and executable model of Spark Streaming on Apache YARN

  • Jan 01, 2020
  • International Journal of Grid and Utility Computing
  • Jia Chun Lin +3
  • PDF
  • Research Article
  • Citations9

Sandbox security model for Hadoop file system

  • Sep 30, 2020
  • Journal of Big Data
  • Gousiya Begum +2
  • Research Article
  • Citations11

Improving Hadoop MapReduce performance on heterogeneous single board computer clusters

  • Jun 15, 2024
  • Future Generation Computer Systems
  • Sooyoung Lim +1
  • Book Chapter
  • Citations3

Interference-aware Workload Scheduling in Co-located Data Centers

  • Jan 01, 2022
  • Ting Zhang +5
  • Research Article
  • Citations86

Thermal-Aware Scheduling of Batch Jobs in Geographically Distributed Data Centers

  • Jan 01, 2014
  • IEEE Transactions on Cloud Computing
  • Marco Polverini +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.