• Home
  • Search
  • Accelerating DNN Inference by Edge-Cloud Collaboration
  • Cite Icon5
  • https://doi.org/10.1109/ipccc51483.2021.9679434Copy DOI Icon

Accelerating DNN Inference by Edge-Cloud Collaboration

  • Oct 29, 2021
  • Jianan Chen +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Deep neural networks (DNN) have become indispensable tools for intelligent applications today. The demand for deploying DNN on the edge devices increases dramatically. Unfortunately, it is challenging because the DNN inference is computation-intensive, but edge devices are always resource-constraint. Prior solutions attempted to address these challenges with collaboration between cloud and edge devices, but they do not take the inference request rate into account. However, the inference delay will increase dramatically while the request rate becomes higher.In this paper, we propose a scheme to dynamic partition DNN into two or three parts and distribute them at the edge and cloud, achieving the lowest delay with the change of request rate. The scheme selects the optimal partition points of DNN with a layer evaluation model (LEM) and a total delay prediction model (DPM) under different request rates. The experiments of distributed deploying AlexNet, VGG, NiN and ResNet DNN models on image classification dataset ImageNet show that the proposed scheme significantly reduces the total end-to-end latency by fully using both the edge and cloud resources. It reduces the inference delay by 1.3 to 1.6 times and improves the throughput 1.2 to 1.7 times compared to the state of art partition approach.

Similar Papers
  • Research Article

Real-time, Work-conserving GPU Scheduling for Concurrent DNN Inference

  • Sep 18, 2025
  • ACM Transactions on Computer Systems
  • Mingcong Han +5
  • Research Article
  • Citations30

Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital-Twin-Assisted Approach

  • Apr 01, 2024
  • IEEE Internet of Things Journal
  • Shisheng Hu +4
  • Conference Article

Accuracy Improvement Methods to a Deep Neural Network Model in Computer Vision

  • May 10, 2022
  • Grace Chrysilla
  • Conference Article

IasRT: Interference-Aware and SLO-Driven GPU Scheduling for Real-Time DNN Inference

  • Nov 10, 2025
  • Heming Zhong +4
  • Conference Article
  • Citations3

An Efficient Event-driven Neuromorphic Architecture for Deep Spiking Neural Networks

  • Sep 01, 2019
  • Duy-Anh Nguyen +3
  • Research Article
  • Citations2

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Ziyang Zhang +4
  • Conference Article
  • Citations4

Capella: Customizing Perception for Edge Devices by Efficiently Allocating FPGAs to DNNs

  • Sep 01, 2019
  • Younmin Bae +4
  • Conference Article
  • Citations1

Resource Constrained Hardware Architecture for Training Deep Neural Networks at the Edge - FPGA Implementation

  • Dec 01, 2022
  • Alavala Venkata Suraj +2
  • Conference Article
  • Citations4

Subthreshold operation of SONOS analog memory to enable accurate low-power neural network inference

  • Dec 03, 2022
  • V Agrawal +14
  • Conference Article
  • Citations9

Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPU

  • Jun 23, 2024
  • Binqi Sun +3
  • Dissertation

Software-hardware co-design: towards ultimate efficiency in deep learning acceleration

  • Jan 01, 2024
  • Peiyan Dong
  • Research Article

Phoenix: Thermal-Aware On-Device Inference of Multi-Instance DNNs for Mobile Video Applications

  • Feb 05, 2026
  • ACM Transactions on Embedded Computing Systems
  • Seunghyeok Jeon +3
  • PDF
  • Research Article
  • Citations10

An Adaptive Task Migration Scheduling Approach for Edge‐Cloud Collaborative Inference

  • Jan 01, 2022
  • Wireless Communications and Mobile Computing
  • Boyin Zhang +4
  • Conference Article
  • Citations24

Adaptive DNN Partition in Edge Computing Environments

  • Dec 01, 2020
  • Weiwei Miao +5
  • Research Article
  • Citations87

Achieving Super-Linear Speedup across Multi-FPGA for Real-Time DNN Inference

  • Oct 08, 2019
  • ACM Transactions on Embedded Computing Systems
  • Weiwen Jiang +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.