• Home
  • Search
  • Performance Analysis of Java Virtual Machine for Machine Learning Workloads using Apache Spark
  • https://doi.org/10.1145/2980258.2982117Copy DOI Icon

Performance Analysis of Java Virtual Machine for Machine Learning Workloads using Apache Spark

  • Aug 25, 2016
  • N Hema +6 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Now a day's data is growing very rapidly, where processing and analyzing data to get useful information is the main task. There are many big data processing tools and framework such as Hadoop, Hive, Cassandra etc. Spark is one of the fastest big data processing framework in cluster computation.Basic Idea is to analyze the performance of java virtual machine (JVM) [1], by characterizing java virtual machine using SparkBench benchmark on Apache Spark™ [2]. Java virtual machine is a core execution platform for spark application. When we run the spark application on java virtual machine, its behavior is affected, which needs to be monitored to analyze the JVM performance. Here we are considering Machine Learning workloads like K-Means, Matrix Factorization and Logistic Regression. Main goal here is to analyze the machine learning workloads end to end across the cluster, with respect to following parameters such as garbage collection, memory such as heap usage, CPU process time. Characterization of JVM is done with spark cluster setup and HDFS is used as storage with distributed Hadoop cluster setup.

Similar Papers
  • Conference Article
  • Citations50

Don't get caught in the cold, warm-up your JVM: understand and eliminate JVM warm-up overhead in data-parallel systems

  • Nov 02, 2016
  • David Lion +5
  • Research Article
  • Citations4

The Impact of Big Data Processing Framework for Artificial Intelligence within Corporate Marketing Communication

  • Dec 13, 2018
  • International Journal of Engineering & Technology
  • Muhamad Fazil Ahmad
  • Conference Article
  • Citations1

Big Data Processing Tools Navigation Diagram

  • Jan 01, 2020
  • Martin Macak +4
  • Conference Article
  • Citations2

A Complex Task Scheduling Scheme for Big Data Platforms Based on Boolean Satisfiability Problem

  • Jul 01, 2018
  • Huang Hong +4
  • Conference Article
  • Citations4

Performance Study of Distributed Big Data Analysis in YARN Cluster

  • Oct 01, 2018
  • Hoo Young Ahn +2
  • Research Article
  • Citations24

VLocality: Revisiting Data Locality for MapReduce in Virtualized Clouds

  • Jan 01, 2017
  • IEEE Network
  • Xiaoqiang Ma +4
  • Research Article
  • Citations7

A configurable and executable model of Spark Streaming on Apache YARN

  • Jan 01, 2020
  • International Journal of Grid and Utility Computing
  • Jia Chun Lin +3
  • PDF
  • Research Article
  • Citations9

Sandbox security model for Hadoop file system

  • Sep 30, 2020
  • Journal of Big Data
  • Gousiya Begum +2
  • PDF
  • Research Article
  • Citations9

Distributed fuzzy clustering algorithm for mixed-mode data in Apache SPARK

  • Dec 21, 2022
  • Journal of Big Data
  • Abdul Wahab Akram +1
  • Research Article
  • Citations11

Improving Hadoop MapReduce performance on heterogeneous single board computer clusters

  • Jun 15, 2024
  • Future Generation Computer Systems
  • Sooyoung Lim +1
  • Book Chapter
  • Citations4

Securing Big Data in Hadoop Using Hybrid Encryption

  • Oct 09, 2021
  • Aswathi Sunder +2
  • Conference Article

Thread-aware garbage collection for server applications

  • Aug 24, 2004
  • Woo Jin Kim +4
  • Conference Article
  • Citations1

Steal-A-GC: Framework to Trigger GC during Idle Periods in Distributed Systems

  • Dec 01, 2016
  • Sujoy Saraswati +2
  • Conference Article
  • Citations8

ALMA

  • Nov 28, 2016
  • Rodrigo Bruno +1
  • Conference Article
  • Citations4

Towards an Efficient Pauseless Java GC with Selective HTM-Based Access Barriers

  • Sep 27, 2017
  • Maria Carpen-Amarie +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.