• Home
  • Search
  • Real-time, Work-conserving GPU Scheduling for Concurrent DNN Inference
  • https://doi.org/10.1145/3768622Copy DOI Icon

Real-time, Work-conserving GPU Scheduling for Concurrent DNN Inference

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Many intelligent applications, such as autonomous driving and virtual reality, require running both latency-critical (real-time) and best-effort deep neural network (DNN) inference tasks to achieve both real-time and work-conserving on the GPU. However, commodity GPUs lack efficient preemptive scheduling support, and existing state-of-the-art approaches either have to monopolize GPU or let real-time tasks to wait for best-effort tasks to complete, resulting in low utilization, high latency, or both. This article presents Reef , the first GPU-accelerated DNN inference serving system that achieves low-latency and work-conserving for concurrent real-time and best-effort tasks. Reef accomplishes this by enabling microsecond-scale kernel preemption and controlled concurrent execution in GPU scheduling. Reef is novel in two ways. First, based on the observation that DNN inference kernels are mostly idempotent, Reef devises a reset-based preemption scheme that launches a real-time kernel on the GPU by proactively killing and restoring best-effort kernels at microsecond-scale. Second, since DNN inference kernels have varied parallelism and predictable latency, Reef proposes a dynamic kernel padding mechanism that dynamically pads the real-time kernel with appropriate best-effort kernels to fully utilize the GPU with negligible overhead. Evaluation using a new DNN inference serving benchmark (DISB) with diverse workloads and a real-world trace on both NVIDIA and AMD GPUs shows that Reef only incurs less than 5% overhead in end-to-end latency for real-time tasks but increases the overall throughput by up to 1.53×, compared to scheduling tasks sequentially. To demonstrate the practical benefits of our approach, we compare Reef with Triton, a widely-adopted production-level serving system. Our evaluation shows that Reef outperforms Triton by 1.12× to 5.20× in end-to-end latency for real-time tasks, while maintaining comparable throughput.

Similar Papers
  • Research Article
  • Citations30

Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital-Twin-Assisted Approach

  • Apr 01, 2024
  • IEEE Internet of Things Journal
  • Shisheng Hu +4
  • Conference Article
  • Citations9

Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPU

  • Jun 23, 2024
  • Binqi Sun +3
  • Conference Article

Accuracy Improvement Methods to a Deep Neural Network Model in Computer Vision

  • May 10, 2022
  • Grace Chrysilla
  • Conference Article

IasRT: Interference-Aware and SLO-Driven GPU Scheduling for Real-Time DNN Inference

  • Nov 10, 2025
  • Heming Zhong +4
  • Conference Article
  • Citations4

Subthreshold operation of SONOS analog memory to enable accurate low-power neural network inference

  • Dec 03, 2022
  • V Agrawal +14
  • Research Article
  • Citations36

ERIDANUS: Efficiently Running Inference of DNNs Using Systolic Arrays

  • Sep 01, 2019
  • IEEE Micro
  • Bahar Asgari +3
  • PDF
  • Research Article
  • Citations2

BestOf: an online implementation selector for the training and inference of deep neural networks

  • May 20, 2022
  • The Journal of Supercomputing
  • Sergio Barrachina +3
  • Research Article
  • Citations87

Achieving Super-Linear Speedup across Multi-FPGA for Real-Time DNN Inference

  • Oct 08, 2019
  • ACM Transactions on Embedded Computing Systems
  • Weiwen Jiang +6
  • Conference Article
  • Citations7

Implementation of Systolic Co-processor for Deep Neural Network Inference based on SoC

  • Nov 01, 2018
  • Erwin Setiawan +1
  • Conference Article
  • Citations6

LBFP: Logarithmic Block Floating Point Arithmetic for Deep Neural Networks

  • Dec 08, 2020
  • Chao Ni +3
  • Preprint Article

ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators

  • Feb 12, 2026
  • arXiv (Cornell University)
  • T Baldi +2
  • Research Article

Phoenix: Thermal-Aware On-Device Inference of Multi-Instance DNNs for Mobile Video Applications

  • Feb 05, 2026
  • ACM Transactions on Embedded Computing Systems
  • Seunghyeok Jeon +3
  • Research Article
  • Citations96

Vega: A Ten-Core SoC for IoT Endnodes With DNN Acceleration and Cognitive Wake-Up From MRAM-Based State-Retentive Sleep Mode

  • Jan 01, 2022
  • IEEE Journal of Solid-State Circuits
  • Davide Rossi +11
  • Conference Article
  • Citations5

Accelerating DNN Inference by Edge-Cloud Collaboration

  • Oct 29, 2021
  • Jianan Chen +4
  • Conference Article
  • Citations7

LoADPart: Load-Aware Dynamic Partition of Deep Neural Networks for Edge Offloading

  • Jul 01, 2022
  • Hongzhou Liu +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.