• Home
  • Search
  • Enabling Serverless Deployment of Large-Scale AI Workloads
  • Cite Icon37
  • https://doi.org/10.1109/access.2020.2985282Copy DOI Icon

Enabling Serverless Deployment of Large-Scale AI Workloads

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We propose a set of optimization techniques for transforming a generic AI codebase so that it can be successfully deployed to a restricted serverless environment, without compromising capability or performance. These involve (1) slimming the libraries and frameworks (e.g., pytorch) used, down to pieces pertaining to the solution; (2) dynamically loading pre-trained AI/ML models into local temporary storage, during serverless function invocation; (3) using separate frameworks for training and inference, with ONNX model formatting; and, (4) performance-oriented tuning for data storage and lookup. The techniques are illustrated via worked examples that have been deployed live on geospatial data from the transportation domain. This draws upon a real-world case study in intelligent transportation looking at on-demand, realtime predictions of flows of train movements across the UK rail network. Evaluation of the proposed techniques shows the response time, for varying volumes of queries involving prediction, to remain almost constant (at 50 ms), even as the database scales up to the 250M entries. The query response time is important in this context as the target is predicting train delays. It is even more important in a serverless environment due to the stringent constraints on serverless functions’ runtime before timeout. The similarities of a serverless environment to other resource constrained environments (e.g., IoT, telecoms) means the techniques can be applied to a range of use cases.

Loading PDF

Similar Papers
  • Conference Article
  • Citations22

Serving Machine Learning Workloads in Resource Constrained Environments: a Serverless Deployment Example

  • Nov 01, 2019
  • Angelos Christidis +2
  • Research Article
  • Citations2

A Bayesian approach to performance modelling for multi-tenant applications using Gaussian models

  • Jan 01, 2016
  • International Journal of High Performance Computing and Networking
  • Junling Zhang +4
  • Single Report
  • Citations40

An efficient compression scheme for bitmap indices

  • Apr 13, 2004
  • Kesheng Wu +2
  • Research Article
  • Citations10

A Comprehensive Risk Management Approach to Information Security in Intelligent Transport Systems

  • May 05, 2021
  • SAE International Journal of Transportation Cybersecurity and Privacy
  • Tom Vogt +13
  • Research Article
  • Citations4

Performance modeling and evaluating workflow of ITS: real-time positioning and route planning

  • Nov 09, 2017
  • Multimedia Tools and Applications
  • Ping Liu +3
  • PDF
  • Research Article
  • Citations58

Emerging smart city, transport and energy trends in urban settings: Results of a pan-European foresight exercise with 120 experts

  • Aug 05, 2022
  • Technological Forecasting and Social Change
  • M Angelidou +4
  • Book Chapter
  • Citations15

Toward a Honeypot Solution for Proactive Security in Vehicular Ad Hoc Networks

  • Jan 01, 2014
  • Dhavy Gantsou +1
  • Book Chapter
  • Citations1

Study of Meta-Data Enrichment Methods to Achieve Near Real Time ETL

  • Nov 05, 2018
  • N Mohammed Muddasir +1
  • Book Chapter
  • Citations2

Chapter 3 - An Agent Methodology for Processes, the Environment, and Services

  • Jan 01, 2015
  • Advances in Artificial Transportation Systems and Simulation
  • Lúcio Sanchez Passos +2
  • Research Article
  • Citations37

Query optimization for parallel execution

  • Jun 01, 1992
  • ACM SIGMOD Record
  • Sumit Ganguly +2
  • Research Article
  • Citations16

Materialized View Selection Using Set Based Particle Swarm Optimization

  • Jul 01, 2018
  • International Journal of Cognitive Informatics and Natural Intelligence
  • Amit Kumar +1
  • Research Article
  • Citations66

Skyframe: a framework for skyline query processing in peer-to-peer systems

  • Jun 24, 2008
  • The VLDB Journal
  • Shiyuan Wang +4
  • Conference Article
  • Citations76

DBEst

  • Jun 25, 2019
  • Qingzhi Ma +1
  • Book Chapter
  • Citations66

Quasi-copies: Efficient data sharing for information retrieval systems

  • Jan 01, 1988
  • Rafael Alonso +3
  • Conference Article
  • Citations5

Flux: Decoupled Auto-Scaling for Heterogeneous Query Workload in Alibaba AnalyticDB

  • Jun 09, 2024
  • Wei Li +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.