• Home
  • Search
  • A Gradient-Aware Search Algorithm for Constrained Markov Decision Processes.
  • Open Access IconOpen Access
  • Cite Icon3
  • https://doi.org/10.1109/tnnls.2023.3315598Copy DOI Icon

A Gradient-Aware Search Algorithm for Constrained Markov Decision Processes.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The canonical solution methodology for finite constrained Markov decision processes (CMDPs), where the objective is to maximize the expected infinite-horizon discounted rewards subject to the expected infinite-horizon discounted costs' constraints, is based on convex linear programming (LP). In this brief, we first prove that the optimization objective in the dual linear program of a finite CMDP is a piecewise linear convex (PWLC) function with respect to the Lagrange penalty multipliers. Next, we propose a novel, provably optimal, two-level gradient-aware search (GAS) algorithm which exploits the PWLC structure to find the optimal state-value function and Lagrange penalty multipliers of a finite CMDP. The proposed algorithm is applied in two stochastic control problems with constraints for performance comparison with binary search (BS), Lagrangian primal-dual optimization (PDO), and LP. Compared with the benchmark algorithms, it is shown that the proposed GAS algorithm converges to the optimal solution quickly without any hyperparameter tuning. In addition, the convergence speed of the proposed algorithm is not sensitive to the initialization of the Lagrange multipliers.

Similar Papers
  • Research Article
  • Citations43

Constrained Soft Actor-Critic for Energy-Aware Trajectory Design in UAV-Aided IoT Networks

  • Jul 01, 2022
  • IEEE Wireless Communications Letters
  • Xuanhan Zhou +4
  • Research Article
  • Citations4

Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints

  • Jun 26, 2023
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Yuhao Ding +1
  • Research Article
  • Citations47

A projection method for the uncapacitated facility location problem

  • Jan 01, 1990
  • Mathematical Programming
  • A R Conn +1
  • Research Article
  • Citations7

Dynamic principal agent model based on CMDP

  • Sep 01, 2003
  • Mathematical Methods of Operations Research (ZOR)
  • Yuanyao Ding +2
  • Research Article
  • Citations12

AoI Minimization Scheme for Short-Packet Communications in Energy-Constrained IIoT

  • Nov 15, 2023
  • IEEE Internet of Things Journal
  • Baoquan Yu +3
  • Research Article
  • Citations185

A novel linear programming approach to fluence map optimization for intensity modulated radiation therapy treatment planning

  • Oct 10, 2003
  • Physics in Medicine & Biology
  • H Edwin Romeijn +4
  • Research Article

Sleeping experts and bandits approach to constrained Markov decision processes

  • Nov 11, 2015
  • Automatica
  • Hyeong Soo Chang
  • Conference Article
  • Citations5

A Systematic Comparative Study of Linear, Binary and Interpolation Search Algorithms

  • Oct 25, 2021
  • Andi Irmayana +4
  • Research Article

Stochastic Power Saving for Macrocell-Assisted Small Cell Networks

  • Apr 26, 2020
  • Journal of Signal Processing Systems
  • Che-Ying Lin +1
  • Research Article
  • Citations6

Delay optimal for reliability-guaranteed concurrent transmissions with raptor code in multi-access 6G edge network

  • Mar 21, 2023
  • Computer Networks
  • Zhongfu Guo +6
  • Conference Article
  • Citations113

Optimal Dynamic Spectrum Access via Periodic Channel Sensing

  • Jan 01, 2007
  • Qianchuan Zhao +3
  • Research Article
  • Citations1

Popularity-Aware Service Provisioning Framework in Cloud Environment

  • Sep 12, 2024
  • Applied Sciences
  • Haneul Ko +3
  • Research Article
  • Citations4

Parameterized Algorithms for MILPs with Small Treedepth

  • May 18, 2021
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Cornelius Brand +2
  • Research Article
  • Citations10

Age of Information Minimization for UAV-Assisted Internet of Things Networks: A Safe Actor-Critic With Policy Distillation Approach

  • Jan 01, 2024
  • IEEE Transactions on Network Science and Engineering
  • Fang Fu +7
  • Research Article
  • Citations33

Convex analysis treated by linear programming

  • Dec 01, 1973
  • Mathematical Programming
  • R J Duffin
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.