• Home
  • Search
  • ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation Models
  • Cite Icon11
  • https://doi.org/10.1109/tmc.2025.3553201Copy DOI Icon

ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation Models

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The prevalence of the pre-training & fine-tuning paradigm enables machine learning models to quickly adapt to various downstream tasks by fine-tuning pre-trained foundation models (FMs), greatly facilitating various IoT applications that rely on model inference in dynamic edge serverless environments. Efficiently dispatching inference requests and configuring instances to batch inference requests can significantly enhance resource efficiency. However, existing serverless inference solutions are tailored for traditional models, make coarse-grained request dispatching and instance configuration decisions, fail to exploit the shared model backbone characteristics of the FM and capture delayed rewards in dynamic environments, and ignore communication latency between edge sites, resulting in high costs and constraint violations. In this paper, we leverage our insight that fine-grained batch inference requests can effectively exploit the shared model backbone feature of FM to save monetary costs. We propose an algorithm that incorporates deep reinforcement learning (DRL) and expert intervention for fine-grained request dispatching and instance configuration, where the DRL component outputs fractional solutions as guidance, while the expert intervention module integrates our insights—batching reduces monetary costs at the expense of increased inference latency, whereas higher configurations shorten inference latency. This module rounds fractional solutions and adjusts instance configurations to search for optimal solutions while satisfying constraints, with theoretical guarantees rigorously proved. Finally, we conducted our experiments on an OpenFaas-based platform and simulator, and extensive trace-driven evaluation results show that ExpertDRL can save costs by up to 85.14% and improve request acceptance ratio by up to 26.93%, compared to the state-of-the-art solution.

Similar Papers
  • Preprint Article

Segmentation Model Benchmarking: A Strategic Prerequisite for Robust Geospatial Foundation Models

  • Mar 14, 2026
  • Mehran Alizadeh Pirbasti +2
  • Research Article

Abstract 5084: Evaluation of single-cell foundation models for cancer outcome predictions

  • Apr 21, 2025
  • Cancer Research
  • Haitham Elmarakeby +3
  • Research Article
  • Citations4

Mixture of Experts-Enabled Parallel Scheduling and Processing for Vehicular Generative AI Services

  • Jan 01, 2026
  • IEEE Transactions on Cognitive Communications and Networking
  • Gaochang Xie +6
  • Dissertation
  • Citations1

Deep reinforcement learning-based dynamic scheduling

  • Jan 01, 2022
  • Renke Liu
  • Research Article
  • Citations1

Deep Reinforcement Learning for irrigation optimization: Advantages, opportunities, and challenges

  • Dec 01, 2025
  • Agricultural Water Management
  • Jiamei Liu +8
  • Dissertation

Location-Aware Application Deployment in Multi-Cloud

  • Jun 20, 2022
  • Tao Shi
  • Research Article

EdgeManager: Online Adaptive Resource Management for Hierarchical DNN Inference in Collaborative Edge Environments

  • Jan 01, 2025
  • IEEE Transactions on Mobile Computing
  • Fengyi Huang +7
  • Research Article
  • Citations27

Hierarchical Control Framework for Path Planning of Mobile Robots in Dynamic Environments Through Global Guidance and Reinforcement Learning

  • Jan 01, 2024
  • IEEE Internet of Things Journal
  • Hongyang Zhao +4
  • Research Article
  • Citations19

FLDQN: Cooperative Multi-Agent Federated Reinforcement Learning for Solving Travel Time Minimization Problems in Dynamic Environments Using SUMO Simulation.

  • Feb 03, 2025
  • Sensors (Basel, Switzerland)
  • Abdul Wahab Mamond +4
  • PDF
  • Research Article
  • Citations38

Deep Reinforcement Learning Unleashing the Power of AI in Decision-Making

  • Feb 02, 2024
  • Journal of Artificial Intelligence General science (JAIGS) ISSN:3006-4023
  • Jeff Shuford
  • Conference Article
  • Citations1

Characterizing Attacks on Deep Reinforcement Learning

  • May 09, 2022
  • Xinlei Pan +8
  • Research Article

O-321 A foundation model for embryology trained on time-lapse images

  • Jul 03, 2024
  • Human Reproduction
  • S Rajendran +8
  • Research Article
  • Citations6

Scalable AP Clustering With Deep Reinforcement Learning for Cell-Free Massive MIMO

  • Jan 01, 2025
  • IEEE Open Journal of the Communications Society
  • Yu Tsukamoto +5
  • Research Article

Research on Dependency-Aware Service Migration Strategy in the Internet of Vehicles Integrating a Graph Attention Network and Deep Reinforcement Learning

  • Jan 16, 2026
  • Applied Sciences
  • Ying Liu +2
  • Research Article
  • Citations7

Integrated Sensing and Communications for UAV Assisted Internet of Things Based on Deep Reinforcement Learning

  • Jun 01, 2025
  • IEEE Transactions on Vehicular Technology
  • Xin Liu +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.