• Home
  • Search
  • Simple and Automatic Distributed Machine Learning on Ray
  • Open Access IconOpen Access
  • Cite Icon1
  • https://doi.org/10.1145/3447548.3470816Copy DOI Icon

Simple and Automatic Distributed Machine Learning on Ray

  • Aug 14, 2021
  • Hao Zhang +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In recent years, the pace of innovations in the fields of machine learning (ML) has accelerated, researchers in SysML have created algorithms and systems that parallelize ML training over multiple devices or computational nodes. As ML models become more structurally complex, many systems have struggled to provide all-round performance on a variety of models. Particularly, ML scale-up is usually underestimated in terms of the amount of knowledge and time required to map from an appropriate distribution strategy to the model. Applying parallel training systems to complex models adds nontrivial development overheads in addition to model prototyping, and often results in lower-than-expected performance. This tutorial identifies research and practical pain points in parallel ML training, and discusses latest development of algorithms and systems on addressing these challenges in both usability and performance. In particular, this tutorial presents a new perspective of unifying seemingly different distributed ML training strategies. Based on it, introduces new techniques and system architectures to simplify and automate ML parallelization. This tutorial is built upon the authors' years' of research and industry experience, comprehensive literature survey, and several latest tutorials and papers published by the authors and peer researchers. The tutorial consists of four parts. The first part will present a landscape of distributed ML training techniques and systems, and highlight the major difficulties faced by real users when writing distributed ML code with big model or big data. The second part dives deep to explain the mainstream training strategies, guided with real use case. By developing a new and unified formulation to represent the seemingly different data- and model- parallel strategies, we describe a set of techniques and algorithms to achieve ML auto-parallelization, and compiler system architectures for auto-generating and exercising parallelization strategies based on models and clusters. The third part of this tutorial exposes a hidden layer of practical pain points in distributed ML training: hyper-parameter tuning and resource allocation, and introduces techniques to improve these aspects. The fourth part is designed as a hands-on coding session, in which we will walk through the audiences on writing distributed training programs in Python, using the various distributed ML tools and interfaces provided by the Ray ecosystem.

Similar Papers
  • Supplementary Content
  • Citations2

Machine Learning Parallelism Could Be Adaptive, Composable and Automated

  • Apr 14, 2021
  • Figshare
  • Hao Zhang
  • Research Article
  • Citations1

Do You Consent to the Use of Your Biological Data for Training ML and AI Models? Online Survey Targeting Clinicians and Researchers.

  • Jan 27, 2024
  • Web3 Journal: ML in Health Science
  • Yury Rusinovich +1
  • Research Article
  • Citations8

Advanced Machine Learning Based Malware Detection Systems

  • Jan 01, 2024
  • IEEE Access
  • Song-Kyoo Kim +5
  • PDF
  • Research Article
  • Citations27

Development of Monthly Reference Evapotranspiration Machine Learning Models and Mapping of Pakistan—A Comparative Study

  • May 23, 2022
  • Water
  • Jizhang Wang +8
  • Research Article

Physics-Informed Machine Learning for Battery Capacity Forecasting

  • Aug 09, 2024
  • Electrochemical Society Meeting Abstracts
  • Mohammad Behtash +5
  • Research Article
  • Citations8

Security in 6G-Based Autonomous Vehicular Networks: Detecting Network Anomalies With Decentralized Federated Learning

  • Mar 01, 2025
  • IEEE Vehicular Technology Magazine
  • Jiazhen Zhang +3
  • Research Article
  • Citations10

Amalur: The Convergence of Data Integration and Machine Learning

  • Dec 01, 2024
  • IEEE Transactions on Knowledge and Data Engineering
  • Ziyu Li +6
  • Preprint Article

The emergence of Continental Crust Revealed by Machine Learning

  • May 15, 2023
  • Chuntao Liu +3
  • Research Article
  • Citations14

HADA: An automated tool for hardware dimensioning of AI applications

  • Jun 11, 2022
  • Knowledge-Based Systems
  • Allegra De Filippo +3
  • Research Article
  • Citations20

Federated learning enables privacy-preserving and data-efficient dimension prediction and part qualification across additive manufacturing factories

  • May 11, 2024
  • Journal of Manufacturing Systems
  • Manan Mehta +4
  • Research Article
  • Citations14

Surrogate-assisted hydraulic fracture optimization workflow with applications for shale gas reservoir development: a comparative study of machine learning models

  • Apr 19, 2022
  • Natural Gas Industry B
  • Cong Xiao +4
  • Research Article
  • Citations26

Deep feature learning of in-cylinder flow fields to analyze cycle-to-cycle variations in an SI engine

  • Dec 04, 2020
  • International Journal of Engine Research
  • Daniel Dreher +9
  • Research Article
  • Citations72

Understanding Software-2.0

  • Jul 23, 2021
  • ACM Transactions on Software Engineering and Methodology
  • Malinda Dilhara +2
  • Conference Article
  • Citations1

Privacy-Preserving in Machine Learning: Differential Privacy Case Study

  • Jan 01, 2024
  • Aleksa Iričanin +2
  • Research Article
  • Citations18

MLLess: Achieving cost efficiency in serverless machine learning training

  • Sep 09, 2023
  • Journal of Parallel and Distributed Computing
  • Pablo Gimeno Sarroca +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.