• Home
  • Search
  • Work-in-Progress: Towards Efficient and Scalable Big Data Analytics: Mapreduce vs. RDD’s
  • Cite Icon4
  • https://doi.org/10.1109/icit.2017.54Copy DOI Icon

Work-in-Progress: Towards Efficient and Scalable Big Data Analytics: Mapreduce vs. RDD’s

  • Dec 1, 2017
  • Sheetanshu Srivastava +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Large-scale data analytics in the recent years has gained prominence due to explosion in data generated from the introduction of varied digital sources such as the internet, social media etc. With the ever increasing size of data and introduction of machine learning methodologies to learn consecutively from this data, distributed programming models such as MapReduce and its open source implementation Hadoop are facing performance issues. The present work discusses the concept of RDDs (Resilient distributed datasets) and it's open source implementation: Apache Spark. Hadoop and MapReduce pay a significant cost for reloading same data from the disk thus wasting significant input/output cycle as well as network bandwidth. Comparisons have been shown in terms of memory occupancy, execution time and lines of codes required in both the paradigms.

Similar Papers
  • Conference Article
  • Citations24

SWAT

  • May 31, 2016
  • Max Grossman +1
  • Research Article
  • Citations12

A design framework for real-time embedded systems with code size and energy constraints

  • Jan 29, 2008
  • ACM Transactions on Embedded Computing Systems
  • Sheayun Lee +4
  • Conference Article
  • Citations1

Distributed information gain theoretic feature selector using spark

  • Dec 01, 2016
  • Bakshi Rohit Prasad +2
  • Research Article

A Hybrid Machine Learning Model for Predictive Analytics in Big Data Frameworks

  • Mar 05, 2025
  • AVE Trends in Intelligent Computing Systems
  • Anjan Kumar Reddy Ayyadapu
  • Supplementary Content

Performance Evaluation of Query Plan Recommendation with Apache Hadoop and Apache Spark

  • Sep 17, 2022
  • arXiv (Cornell University)
  • Elham Azhir +3
  • PDF
  • Research Article
  • Citations116

Sehaa: A Big Data Analytics Tool for Healthcare Symptoms and Diseases Detection Using Twitter, Apache Spark, and Machine Learning

  • Feb 19, 2020
  • Applied Sciences
  • Shoayee Alotaibi +4
  • Book Chapter
  • Citations5

Using Apache Spark

  • Jan 01, 2016
  • Deepak Vohra
  • Research Article

How Does Software Configuration Parameters Impact Job's Execution Time in Spark?

  • May 22, 2025
  • Journal of Internet Services and Applications
  • Maria C L Nunes +1
  • Book Chapter
  • Citations2

Fake News Detection Approach Using Parallel Predictive Models and Spark to Avoid Misinformation Related to Covid-19 Epidemic

  • Jan 01, 2021
  • Youness Madani +2
  • Research Article
  • Citations7

Cloud computing and big data: Technologies and applications

  • May 20, 2018
  • Concurrency and Computation: Practice and Experience
  • Mostapha Zbakh +3
  • Research Article
  • Citations17

ParSoDA: high-level parallel programming for social data mining

  • Dec 19, 2018
  • Social Network Analysis and Mining
  • Loris Belcastro +3
  • PDF
  • Research Article

Novel Apache Spark based Algorithm to Solve Dirichlet Problem for Poisson Equation in 3D Computational Domain

  • Oct 01, 2016
  • Journal of Computer Science
  • Shomanov Aday +1
  • Research Article
  • Citations50

Scheduling Jobs across Geo-Distributed Datacenters with Max-Min Fairness

  • Jul 01, 2019
  • IEEE Transactions on Network Science and Engineering
  • Li Chen +3
  • Research Article
  • Citations10

Secure cloud model for intellectual privacy protection of arithmetic expressions in source codes using data obfuscation techniques

  • Apr 26, 2022
  • Theoretical Computer Science
  • Pallavi Ahire +1
  • Conference Article
  • Citations2

Programming fuzzy logic in assembly language

  • Nov 04, 1996
  • J.M Sibigtroth
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.