• Home
  • Search
  • Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster
  • Cite Icon2
  • https://doi.org/10.1145/3651890.3672241Copy DOI Icon

Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster

  • Aug 4, 2024
  • Xuya Jia +13 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Big data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions.

Similar Papers
  • Conference Article
  • Citations24

Adaptive Failure Detection via Heartbeat under Hadoop

  • Dec 01, 2011
  • Hao Zhu +1
  • Research Article
  • Citations50

Scheduling Jobs across Geo-Distributed Datacenters with Max-Min Fairness

  • Jul 01, 2019
  • IEEE Transactions on Network Science and Engineering
  • Li Chen +3
  • Research Article

Dynamic Multi-objective Resource Optimization in Big Data Clusters

  • Aug 12, 2021
  • International Journal of Innovative Research in Engineering & Multidisciplinary Physical Sciences
  • Kanagalakshmi Murugan
  • Research Article
  • Citations3

The Role of Mobile Communications for Industrial Automation: Architecture, Applications and Challenges

  • Jan 01, 2025
  • IEEE Open Journal of the Communications Society
  • Engin Zeydan +4
  • Research Article
  • Citations3

Is It An Interesting Job, And Will I Persist, Perform, And Be More Content? A Quasi-Experimental Investigation

  • Jan 01, 2017
  • Performance Improvement Quarterly
  • Sacip Toker
  • Conference Article
  • Citations5

Stochastic Scheduling on Unrelated Machines

  • May 11, 2017
  • DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
  • Martin Skutella +2
  • Conference Article

Analysis of Job Completion Time in Vehicular Cloud Under Concurrent Task Execution

  • Feb 20, 2023
  • Chinh Tran +1
  • Research Article
  • Citations4

Optimal job fragmentation

  • Oct 16, 2009
  • ACM SIGMETRICS Performance Evaluation Review
  • Jayakrishnan Nair +1
  • Research Article
  • Citations12

Scheduling no-wait production with time windows and flexible processing times

  • Jan 01, 2001
  • IEEE Transactions on Robotics and Automation
  • F Chauvet +2
  • Research Article
  • Citations1

Online Job Dispatching and Scheduling to Minimize Job Completion Time and to Meet Deadlines

  • Dec 01, 2018
  • Journal of Interconnection Networks
  • Yupeng Li
  • Research Article
  • Citations34

Deadline-Aware MapReduce Job Scheduling with Dynamic Resource Availability

  • Apr 01, 2019
  • IEEE Transactions on Parallel and Distributed Systems
  • Dazhao Cheng +4
  • Conference Article
  • Citations19

SchedTune: A Heterogeneity-Aware GPU Scheduler for Deep Learning

  • May 01, 2022
  • Hadeel Albahar +5
  • Research Article
  • Citations55

Modeling NaTech-related domino effects in process clusters: A network-based approach

  • Jan 25, 2022
  • Reliability Engineering & System Safety
  • Meng Lan +5
  • Book Chapter
  • Citations1

Improved PC Based Resource Scheduling Algorithm for Virtual Machines in Cloud Computing

  • Jan 01, 2016
  • Baiyou Qiao +7
  • Research Article
  • Citations161

Simultaneous Minimization of Mean and Variation of Flow Time and Waiting Time in Single Machine Systems

  • Feb 01, 1989
  • Operations Research
  • Uttarayan Bagchi
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.