• Home
  • Search
  • Surviving failures in bandwidth-constrained datacenters
  • Open Access IconOpen Access
  • Cite Icon36
  • https://doi.org/10.1145/2377677.2377760Copy DOI Icon

Surviving failures in bandwidth-constrained datacenters

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Datacenter networks have been designed to tolerate failures of network equipment and provide sufficient bandwidth. In practice, however, failures and maintenance of networking and power equipment often make tens to thousands of servers unavailable, and network congestion can increase service latency. Unfortunately, there exists an inherent tradeoff between achieving high fault tolerance and reducing bandwidth usage in network core; spreading servers across fault domains improves fault tolerance, but requires additional bandwidth, while deploying servers together reduces bandwidth usage, but also decreases fault tolerance. We present a detailed analysis of a large-scale Web application and its communication patterns. Based on that, we propose and evaluate a novel optimization framework that achieves both high fault tolerance and significantly reduces bandwidth usage in the network core by exploiting the skewness in the observed communication patterns.

Similar Papers
  • Research Article
  • Citations61

Analyzing, modeling and evaluating dynamic adaptive fault tolerance strategies in cloud computing environments

  • Mar 21, 2013
  • The Journal of Supercomputing
  • Dawei Sun +3
  • Conference Article
  • Citations4

Node covering, error correcting codes and multiprocessors with very high average fault tolerance

  • Jun 27, 1995
  • S Dutt +1
  • PDF
  • Research Article
  • Citations23

Balanced Energy-Aware and Fault-Tolerant Data Center Scheduling

  • Feb 14, 2022
  • Sensors (Basel, Switzerland)
  • Muhammad Shaukat +5
  • Conference Article

RingS: A Novel Structured P2P Network Model with Explicit Locality and High Fault Tolerance

  • Nov 01, 2013
  • Zhenhua Wang +4
  • Conference Article
  • Citations4

Design of Fault Tolerant Universal Logic in QCA

  • Dec 01, 2014
  • Bibhash Sen +3
  • Research Article
  • Citations23

Design and Analysis of a New Fault-Tolerant Magnetic-Geared Permanent-Magnet Motor

  • Jun 01, 2014
  • IEEE Transactions on Applied Superconductivity
  • Guohai Liu +4
  • Conference Article
  • Citations2

Image Processing and Deep Learning Technology Help Power Equipment Intelligent Operation Inspection

  • Dec 23, 2022
  • Zhang Shiling +1
  • Conference Article
  • Citations2

A Lightweight Named Entity Recognition Method for Chinese Power Equipment Defect Text

  • Nov 04, 2022
  • Yifan Jiang +3
  • Research Article
  • Citations3

Scaling-up versus scaling-out networking in data centers: a comparative robustness analysis

  • May 08, 2018
  • The Journal of Supercomputing
  • L Shooshtarian +2
  • Conference Article

The design of reliable VLSI electronic systems for digital signal processing

  • Jan 01, 1991
  • W.K Jenkins
  • Research Article
  • Citations4

Triple redundant fault tolerance: A hardware-implemented approach

  • Jan 01, 1991
  • ISA Transactions
  • Steven E Smith
  • Book Chapter
  • Citations1

Toward PTNET Network Topology Analysis and Routing Algorithm Design

  • Jan 01, 2020
  • Zhijie Han +4
  • Research Article
  • Citations2

Integrating Edge Computing And Iot For Real-Time Air And Water Quality Monitoring Systems

  • Jul 17, 2025
  • International Journal of Environmental Sciences
  • H L Yadav +5
  • Book Chapter

Carbon Footprint Quantification Algorithm in the Life Cycle of Power Equipment

  • Jan 01, 2023
  • Chen Shang +8
  • Conference Article
  • Citations1

Thermal Analysis of Asymmetric Winding Multiphase Motor

  • Aug 01, 2019
  • Zhao Tianxu +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.