• Cite Icon24
  • https://doi.org/10.1145/2907294.2907307Copy DOI Icon

SWAT

  • May 31, 2016
  • Max Grossman +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The field of data analytics is currently going through a renaissance as a result of ever-increasing dataset sizes, the value of the models that can be trained from those datasets, and a surge in flexible, distributed programming models. In particular, the Apache Hadoop and Spark programming systems, as well as their supporting projects (e.g. HDFS, SparkSQL), have greatly simplified the analysis and transformation of datasets whose size exceeds the capacity of a single machine. While these programming models facilitate the use of distributed systems to analyze large datasets, they have been plagued by performance issues. The I/O performance bottlenecks of Hadoop are partially responsible for the creation of Spark. Performance bottlenecks in Spark due to the JVM object model, garbage collection, interpreted/managed execution, and other abstraction layers are responsible for the creation of additional optimization layers, such as Project Tungsten. Indeed, the Project Tungsten issue tracker states that the "majority of Spark workloads are not bottlenecked by I/O or network, but rather CPU and memory".

Similar Papers
  • Conference Article
  • Citations4

Work-in-Progress: Towards Efficient and Scalable Big Data Analytics: Mapreduce vs. RDD’s

  • Dec 01, 2017
  • Sheetanshu Srivastava +2
  • Supplementary Content
  • Citations3

Performance evaluation of GPU- and cluster-computing for parallelization of compute-intensive tasks

  • Aug 06, 2021
  • International Journal of Web Information Systems
  • Alexander Döschl +2
  • Research Article

A Hybrid Machine Learning Model for Predictive Analytics in Big Data Frameworks

  • Mar 05, 2025
  • AVE Trends in Intelligent Computing Systems
  • Anjan Kumar Reddy Ayyadapu

Big Data Tools: A Survey

  • Jan 14, 2021
  • Shwetha M Banakar +2
  • Book Chapter
  • Citations5

Using Apache Spark

  • Jan 01, 2016
  • Deepak Vohra
  • Conference Article
  • Citations17

Spark-DIY: A Framework for Interoperable Spark Operations with High Performance Block-Based Data Models

  • Dec 01, 2018
  • Silvina Caino-Lores +4
  • Research Article
  • Citations1

DNA barcoding using particle swarm optimization on apache spark SQL case study: DNA of covid-19

  • Jan 01, 2021
  • International Journal of Nonlinear Analysis and Applications
  • Lala Septem Riza +3
  • Supplementary Content

Performance Evaluation of Query Plan Recommendation with Apache Hadoop and Apache Spark

  • Sep 17, 2022
  • arXiv (Cornell University)
  • Elham Azhir +3
  • Research Article
  • Citations16

Integration of Spark framework in Supply Chain Management

  • Jan 01, 2016
  • Procedia Computer Science
  • Harjeet Singh Jaggi +1
  • Book Chapter
  • Citations58

SPARQLGX: Efficient Distributed Evaluation of SPARQL with Apache Spark

  • Jan 01, 2016
  • Damien Graux +3
  • Conference Article
  • Citations9

MCS: Memory Constraint Strategy for Unified Memory Manager in Spark

  • Dec 01, 2017
  • Ziyao Zhu +3
  • Conference Article
  • Citations18

Understanding and improving disk-based intermediate data caching in Spark

  • Dec 01, 2017
  • Kaihui Zhang +3
  • Conference Article
  • Citations10

Jointly optimizing task granularity and concurrency for in-memory mapreduce frameworks

  • Dec 01, 2017
  • Jonghyun Bae +7
  • Research Article
  • Citations1

Distributed Systems and Big Data Analytics in Predictive Healthcare: Transforming Modern Medicine

  • Mar 25, 2025
  • International Journal of Scientific Research in Computer Science, Engineering and Information Technology
  • Shridhar Bhalekar
  • Research Article

Big Data Analytics for Business Growth: Leveraging Large-Scale Data for Competitive Advantage

  • Dec 09, 2025
  • European Journal of Applied Science, Engineering and Technology
  • Syed Mohammed Walid Karim +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.