• Home
  • Search
  • Achieving Load Balance for Parallel Data Access on Distributed File Systems
  • Cite Icon35
  • https://doi.org/10.1109/tc.2017.2749229Copy DOI Icon

Achieving Load Balance for Parallel Data Access on Distributed File Systems

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The distributed file system, HDFS, is widely deployed as the bedrock for many parallel big data analysis. However, when running multiple parallel applications over the shared file system, the data requests from different processes/executors will unfortunately be served in a surprisingly imbalanced fashion on the distributed storage servers. These imbalanced access patterns among storage nodes are caused because a). unlike conventional parallel file system using striping policies to evenly distribute data among storage nodes, data-intensive file system such as HDFS store each data unit, referred to as chunk file, with several copies based on a relative random policy, which can result in an uneven data distribution among storage nodes; b). based on the data retrieval policy in HDFS, the more data a storage node contains, the higher probability the storage node could be selected to serve the data. Therefore, on the nodes serving multiple chunk files, the data requests from different processes/executors will compete for shared resources such as hard disk head and networkbandwidth, resulting in a degraded I/O performance. In this paper, we first conduct a complete analysis on how remote and imbalanced read/write patterns occur and how they are affected by the size of the cluster. We then propose novel methods, referred to as Opass, to optimize parallel data reads, as well as to reduce the imbalance of parallel writes on distributed file systems. Our proposed methods can benefit parallel data-intensive analysis with various parallel data access strategies. Opass adopts new matching-based algorithms to match processes to data so as to compute the maximum degree of data locality and balanced data access. Furthermore, to reduce the imbalance of parallel writes, Opass employs a heatmap for monitoring the I/O statuses of storage nodes and performs HM-LRU policy to select a local optimal storage node for serving write requests. Experiments are conducted on PRObE's Marmot 128-node cluster testbed and the results from both benchmark and well-known parallel applications show the performance benefits and scalability of Opass.

Similar Papers
  • Book Chapter
  • Citations3

Zebra: An Efficient, RDMA-Enabled Distributed Persistent Memory File System

  • Jan 01, 2022
  • Jingyu Wang +4
  • Book Chapter
  • Citations3

A Distributed File System

  • Feb 27, 2006
  • J S Yadav +1
  • Conference Article
  • Citations6

Evolution and analysis of distributed file systems in cloud storage: Analytical survey

  • Apr 01, 2016
  • Dharavath Ramesh +3
  • Abstract

152. Pain plan implementation decreases postoperative opioid use, hospital length of stay and clinic resource utilization for patients undergoing elective spine surgery

  • Aug 10, 2021
  • The Spine Journal
  • Harjot Uppal +5
  • Conference Article
  • Citations12

Implement a reliable and secure cloud distributed file system

  • Nov 01, 2012
  • Fan-Hsun Tseng +3
  • Conference Article
  • Citations3

Wofs: A Distributed Network File System Supporting Fast Data Insertion and Truncation

  • May 01, 2010
  • Cheng-Chia Wang +1
  • Book Chapter
  • Citations1

Optimization of Massively Parallel Data Flows

  • Nov 28, 2013
  • Fabian Hueske +1
  • Research Article
  • Citations5

A General-Purpose Architecture for Replicated Metadata Services in Distributed File Systems

  • Oct 01, 2017
  • IEEE Transactions on Parallel and Distributed Systems
  • Dimokritos Stamatakis +3
  • Conference Article
  • Citations4

TLDFS: A Distributed File System based on the Layered Structure

  • Sep 01, 2007
  • Lei Wang +1
  • Research Article
  • Citations157

Load Rebalancing for Distributed File Systems in Clouds

  • May 01, 2013
  • IEEE Transactions on Parallel and Distributed Systems
  • Hung-Chang Hsiao +3
  • Book Chapter
  • Citations10

Evaluating delayed write in a multilevel caching file system

  • Jan 01, 1996
  • Daniel A Muntz +2
  • Research Article
  • Citations4

VM/ESA support for coordinated recovery of files

  • Jan 01, 1991
  • IBM Systems Journal
  • C C Barnes +3
  • Conference Article
  • Citations16

An Insight about GlusterFS and Its Enforcement Techniques

  • May 01, 2016
  • Manikandan Selvaganesan +1
  • Conference Article
  • Citations8

The Design and Evaluation of a Distributed Reliable File System

  • Dec 01, 2009
  • Dalibor Peric +4
  • PDF
  • Conference Article
  • Citations2

Application performance on the Direct Access File System

  • Jan 01, 2004
  • Alexandra Fedorova +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.