• Home
  • Search
  • Optimize Large-Scale Data Processing Via Spark Tuning
  • https://doi.org/10.56726/irjmets45567Copy DOI Icon

Optimize Large-Scale Data Processing Via Spark Tuning

  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Apache Spark is a leading open-source data processing engine used for batch processing, machine learning, stream processing, and large-scale SQL (structured query language).It has been designed to make big data processing quicker and easier.Since its inception, Spark has gained huge popularity as a big data processing framework and is extensively used by different industries and businesses that are dealing with large volumes of data.This paper will exhibit actionable solutions to maximize our chances of reducing computation time by optimizing Spark jobs.The strategy lays out different run stages, wherein each run stage builds upon the previous and improves the computation time by making new enhancements and recommendations.

Similar Papers
  • Research Article
  • Citations4

The Impact of Big Data Processing Framework for Artificial Intelligence within Corporate Marketing Communication

  • Dec 13, 2018
  • International Journal of Engineering & Technology
  • Muhamad Fazil Ahmad
  • Conference Article
  • Citations2

A Complex Task Scheduling Scheme for Big Data Platforms Based on Boolean Satisfiability Problem

  • Jul 01, 2018
  • Huang Hong +4
  • Research Article

Big Data Processing and Data Analytics

  • May 22, 2022
  • INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Sheik Mohammed Shaw S
  • Conference Article
  • Citations42

Network security and anomaly detection with Big-DAMA, a big data analytics framework

  • Sep 01, 2017
  • Pedro Casas +4
  • PDF
  • Research Article
  • Citations9

Distributed fuzzy clustering algorithm for mixed-mode data in Apache SPARK

  • Dec 21, 2022
  • Journal of Big Data
  • Abdul Wahab Akram +1
  • Conference Article
  • Citations3

Docker environment based Apache Storm and Spark Benchmark Test

  • Sep 01, 2020
  • Jiwon Bang +1
  • Conference Article
  • Citations4

Performance Study of Distributed Big Data Analysis in YARN Cluster

  • Oct 01, 2018
  • Hoo Young Ahn +2
  • Research Article
  • Citations7

A configurable and executable model of Spark Streaming on Apache YARN

  • Jan 01, 2020
  • International Journal of Grid and Utility Computing
  • Jia Chun Lin +3
  • Research Article
  • Citations7

Cloud computing and big data: Technologies and applications

  • May 20, 2018
  • Concurrency and Computation: Practice and Experience
  • Mostapha Zbakh +3
  • Conference Article

Performance Analysis of Java Virtual Machine for Machine Learning Workloads using Apache Spark

  • Aug 25, 2016
  • N Hema +6
  • Research Article
  • Citations53

In-Mapper combiner based MapReduce algorithm for processing of big climate data

  • Apr 18, 2018
  • Future Generation Computer Systems
  • Gunasekaran Manogaran +2
  • Conference Article
  • Citations7

Cyclone: Unified Stream and Batch Processing

  • Aug 01, 2016
  • Matu Harvan +2
  • Research Article
  • Citations1

Research on Communication Quality Monitoring System Driven by Big Data in C/S Architecture

  • Jan 01, 2024
  • Computing, Performance and Communication Systems
  • Yiru Zhang
  • Research Article

Optimized design of hardware accelerator based on machine learning

  • Aug 19, 2024
  • Highlights in Science, Engineering and Technology
  • Jixuan Liu
  • Conference Article
  • Citations56

Improving spark application throughput via memory aware task co-location

  • Dec 11, 2017
  • Vicent Sanz Marco +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.