• Home
  • Search
  • Resource-Efficient Orchestration for Heterogeneous Serverless Computing with Harmonized Effectiveness and Practicability
  • https://doi.org/10.1145/3788863Copy DOI Icon

Resource-Efficient Orchestration for Heterogeneous Serverless Computing with Harmonized Effectiveness and Practicability

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Current serverless platforms struggle to optimize resource utilization for both CPU and GPU functions due to their dynamic and fine-grained nature. Conventional techniques like overcommitment and autoscaling fall short, often sacrificing utilization for practicability or incurring performance trade-offs. Overcommitment requires predicting performance to prevent QoS violation, introducing trade-off between prediction accuracy and overheads. Autoscaling requires scaling instances in response to load fluctuations quickly to reduce resource wastage, but more frequent scaling also leads to more cold start overheads. The rich concurrency of GPU resources further complicates GPU instance orchestration, such as setting right batch sizes. This paper introduces Jiagu to harmonize efficiency with practicability through the following novel techniques. First, pre-decision scheduling achieves accurate prediction while eliminating overheads by decoupling prediction and scheduling. Second, dual-staged scaling achieves frequent adjustment of instances with minimum overhead. Third, Jiagu conducts an in-depth analysis about the complexity of the relationship between GPU function configuration and execution. It then proposes batch-aware scaling that achieves optimal configurations for both batch size setting and autoscaling, addressing all the challenges according to the analysis. We have implemented a prototype and evaluated it using real-world applications and traces from the public cloud platform. Our evaluation shows an improvement in deployment density over commercial clouds (with Kubernetes) while maintaining QoS for both CPU and GPU functions (54.8% and 18% respectively), and 81.0%–93.7% lower scheduling costs and a 57.4%–69.3% reduction in cold start latency compared to existing QoS-aware schedulers.

Similar Papers
  • Conference Article
  • Citations4

Security, Compliance, and Agile Deployment of Personal Identifiable Information Solutions on a Public Cloud

  • Jun 01, 2016
  • Yasuharu Katsuno +6
  • Preprint Article

Project Pythia: Building an Inclusive Geoscience Community with Cookbooks

  • Mar 11, 2024
  • John Clyne +10
  • Research Article

Modeling user autonomy in recommender systems using Markov perturbation-based multi-armed bandits

  • Feb 21, 2025
  • Theoretical and Natural Science
  • Yuxuan Yang
  • Conference Article
  • Citations5

Ontology integration for advanced manufacturing collaboration in cloud platforms

  • May 01, 2015
  • Shravya Ramisetty +5
  • Conference Article
  • Citations1

Network-based Intrusion Prevention System Prototype with Multi-Detection - A Position Paper

  • Jan 01, 2014
  • Daniel Kavan +2
  • Conference Article
  • Citations43

Game Theoretic Modeling of Security and Interdependency in a Public Cloud

  • Jun 01, 2014
  • Charles A Kamhoua +5
  • Conference Article
  • Citations9

Cross-Layer SLA Management for Cloud-hosted Big Data Analytics Applications

  • May 01, 2015
  • Xuezhi Zeng +4
  • Conference Article
  • Citations32

Detecting performance interference in cloud-based web services

  • May 01, 2015
  • Yasaman Amannejad +2
  • Research Article
  • Citations27

The operational cost minimization in distributed clouds via community-aware user data placements of social networks

  • Nov 23, 2016
  • Computer Networks
  • Qiufen Xia +2
  • Research Article
  • Citations4

Serverless Computing: The Future of Scalability and Efficiency with AWS, Azure, and GCP

  • Feb 26, 2025
  • International Journal of Advanced Research in Science, Communication and Technology
  • Praveen Borra +1
  • PDF
  • Research Article
  • Citations19

Optimization of Edge Resources for Deep Learning Application with Batch and Model Management

  • Sep 05, 2022
  • Sensors (Basel, Switzerland)
  • Seungwoo Kum +3
  • Conference Article
  • Citations37

Tackling Cold Start of Serverless Applications by Efficient and Adaptive Container Runtime Reusing

  • Sep 01, 2021
  • Kun Suo +4
  • Research Article
  • Citations1

Implementation of Gated Recurrent Unit, Long Short-Term Memory and Derivatives for Gold Price Prediction

  • Jan 12, 2025
  • Public Research Journal of Engineering, Data Technology and Computer Science
  • Amanda Iksanul Putri +3
  • Research Article
  • Citations1

Neural networks in catchment hydrology: a comparative study of different algorithms in an ensemble of ungauged basins in Germany

  • Oct 14, 2025
  • Hydrology and earth system sciences
  • Max Weißenborn +2
  • Conference Article
  • Citations12

TomusBlobs: Towards Communication-Efficient Storage for MapReduce Applications in Azure

  • Feb 15, 2012
  • Radu Tudoran +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.