• Home
  • Search
  • Shared Memory Based RDD Data Sharing on Spark
  • Open Access IconOpen Access
  • https://doi.org/10.12783/dtetr/ssme-ist2016/3939Copy DOI Icon

Shared Memory Based RDD Data Sharing on Spark

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Apache Spark is an increasingly popular fast big data analytics engine, which focuses on a large-scale data processing. Currently, Spark's memory management mainly oriented to single application, and do not directly support the typical scenarios of multiple data processing applications. In this paper, we propose an extension of Apache Spark, called Shared Memory Spark (SMSpark). SMSpark introduces shared memory based on RDD data sharing between applications. We conducted experiment in a cluster using two typical applications. Experimental results show that compared to Spark, SMSpark gains better performance.

Similar Papers
  • Conference Article
  • Citations35

SparkSW: Scalable Distributed Computing System for Large-Scale Biological Sequence Alignment

  • May 01, 2015
  • Guoguang Zhao +2
  • Research Article
  • Citations7

Cloud computing and big data: Technologies and applications

  • May 20, 2018
  • Concurrency and Computation: Practice and Experience
  • Mostapha Zbakh +3
  • Research Article
  • Citations41

Mammoth: Gearing Hadoop Towards Memory-Intensive MapReduce Applications

  • Aug 01, 2015
  • IEEE Transactions on Parallel and Distributed Systems
  • Xuanhua Shi +7
  • Conference Article
  • Citations3

Docker environment based Apache Storm and Spark Benchmark Test

  • Sep 01, 2020
  • Jiwon Bang +1
  • Conference Article
  • Citations4

Distributed Large-scale Time-series Data Processing and Analysis System Based on Spark Platform

  • Jun 01, 2021
  • Bangyan Du
  • Conference Article
  • Citations7

Conquering big data with spark and BDAS

  • Jun 16, 2014
  • Ion Stoica
  • Research Article

Optimize Large-Scale Data Processing Via Spark Tuning

  • Oct 27, 2023
  • International Research Journal of Modernization in Engineering Technology and Science
  • Apurva Kumar +1
  • Research Article
  • Citations55

Leveraging the power of multi-core platforms for large-scale geospatial data processing: Exemplified by generating DEM from massive LiDAR point clouds

  • Apr 27, 2010
  • Computers & Geosciences
  • Xuefeng Guan +1
  • Conference Article
  • Citations9

MCS: Memory Constraint Strategy for Unified Memory Manager in Spark

  • Dec 01, 2017
  • Ziyao Zhu +3
  • Dissertation

Multi-constraint scheduling of MapReduce workloads

  • Jul 15, 2014
  • Jordà Polo Bardés
  • Conference Article

Symbiosis in scale out networking and data management

  • Sep 18, 2012
  • Amin Vahdat
  • Research Article
  • Citations8

A Hybrid Mechanism of Particle Swarm Optimization and Differential Evolution Algorithms based on Spark

  • Dec 31, 2019
  • KSII Transactions on Internet and Information Systems
  • Debin Fan +1
  • Book Chapter
  • Citations3

PPCAS: Implementation of a Probabilistic Pairwise Model for Consistency-Based Multiple Alignment in Apache Spark

  • Jan 01, 2017
  • Jordi Lladós +2
  • Conference Article
  • Citations1

Assessing the computational limits of GraphDBs’ engines - A comparison study between Neo4j and Apache Spark

  • Nov 20, 2020
  • Ioannis Ballas +3
  • Research Article

Implementation of a Big Data Ecosystem for Interdisciplinary Analysis in Academic Projects: Collaboration between Engineering and Communication in the Study of the 2024 Mexican Elections

  • Dec 17, 2024
  • Revista de Políticas Universitarias
  • Omar Mendoza-González +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.