• Home
  • Search
  • M-VP2: Microservice-Oriented Vulnerability Patch Planning - A Cost-Aware Approach Using Multi-Agent Reinforcement Learning
  • https://doi.org/10.20944/preprints202601.1784.v1Copy DOI Icon

M-VP2: Microservice-Oriented Vulnerability Patch Planning - A Cost-Aware Approach Using Multi-Agent Reinforcement Learning

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Microservice architectures amplify the volume and complexity of security vulnerabilities, making it increasingly difficult for security and SRE teams to decide which services to patch, when to patch them, and how to coordinate patches under strict cost and availability constraints. Traditional prioritization schemes based on CVSS scores or static business criticality heuristics ignore inter-service dependencies, deployment topologies, and operational costs such as downtime, rollback risk, and engineering effort. In this paper, we propose M-VP2a microservice-oriented vulnerability patch planning framework that formulates patch scheduling as a cost-aware multi-agent reinforcement learning (MARL) problem. Each microservice is modeled as an autonomous agent that selects patching actions over time (e.g. patch now, defer, or batch with other changes), while a joint reward function balances security risk reduction, patching and downtime cost, and compliance with service-level objectives. The environment captures call-graph dependencies, cascading failure modes, and temporal exploit likelihood, enabling agents to learn coordination strategies that avoid risky simultaneous updates on tightly coupled services. We design a hierarchical actor–critic architecture with centralized training and decentralized execution, augmented with a risk-aware reward shaping mechanism to penalize unsafe patch combinations and SLA violations. Extensive simulation experiments on synthetic and real-world–inspired microservice topologies show that M-VP2 reduces expected breach risk and aggregate patching cost by up to double-digit percentages compared with CVSS-based heuristics, greedy risk–cost ranking, and single-agent RL baselines, while producing patch plans that are more stable, interpretable, and aligned with operational constraints.

Similar Papers
  • Conference Article
  • Citations18

Interaction-Aware Multi-Agent Reinforcement Learning for Mobile Agents with Individual Goals

  • May 01, 2019
  • Anahita Mohseni-Kabir +2
  • Research Article

LLM Collaboration with Multi-Agent Reinforcement Learning

  • Mar 14, 2026
  • Shuo Liu +3
  • Conference Article
  • Citations1

Multi-agent Robust Time Differential Reinforcement Learning Over Communicated Networks

  • Jul 01, 2018
  • Jiahong Li +2
  • Conference Article
  • Citations151

Networked Multi-Agent Reinforcement Learning in Continuous Spaces

  • Dec 01, 2018
  • Kaiqing Zhang +2
  • Research Article

Algorithmic Approach to Prescriptive Maintenance in Industry 4.0

  • Mar 01, 2026
  • Production Engineering Archives
  • Piotr Wittbrodt
  • Research Article
  • Citations17

Secrecy Rate Maximization in THz-Aided Heterogeneous Networks: A Deep Reinforcement Learning Approach

  • Oct 01, 2023
  • IEEE Transactions on Vehicular Technology
  • Himanshu Sharma +3
  • Research Article

A Decentralized Actor-Critic Algorithm With Entropy Regularization and Its Finite-Time Analysis.

  • Jan 01, 2025
  • IEEE transactions on neural networks and learning systems
  • Tao Mao +5
  • Conference Article
  • Citations20

Multi - Agent Reinforcement Learning for Spectrum Sharing in Vehicular Networks

  • Jul 01, 2019
  • Le Liang +2
  • PDF
  • Research Article
  • Citations131

Multi-agent reinforcement learning for cooperative lane changing of connected and autonomous vehicles in mixed traffic

  • Mar 16, 2022
  • Autonomous Intelligent Systems
  • Wei Zhou +5
  • Research Article
  • Citations22

A benefit-aware on-demand provisioning approach for multi-tier applications in cloud computing

  • Jul 28, 2013
  • Frontiers of Computer Science
  • Heng Wu +4
  • Research Article
  • Citations6

A Deep Neural Network-Based Multi-Label Classifier for SLA Violation Prediction in a Latency Sensitive NFV Application

  • Jan 01, 2021
  • IEEE Open Journal of the Communications Society
  • Nikita Jalodia +2
  • Research Article

A Multi-Agent Reinforcement Framework for Autonomous Cloud Resource Scheduling and Optimization

  • Jan 01, 2024
  • American International Journal of Computer Science and Technology
  • Helmi Laura
  • Research Article

Robust and efficient communication in multi-agent reinforcement learning.

  • Feb 01, 2026
  • Chaos (Woodbury, N.Y.)
  • Zejiao Liu +8
  • Conference Article
  • Citations4

Combining multi-agent systems and MDE approach for monitoring SLA violations in the Cloud Computing

  • Jun 01, 2015
  • Adil Maarouf +3
  • Conference Article
  • Citations4

On the Use of Traffic Information to Improve the Coordinated P2P Detection of SLA Violations

  • May 01, 2014
  • Jeferson C Nobre +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.