• Home
  • Search
  • Deterministic policy gradient algorithms for semi‐Markov decision processes
  • Cite Icon6
  • https://doi.org/10.1002/int.22709Copy DOI Icon

Deterministic policy gradient algorithms for semi‐Markov decision processes

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

A large class of sequential decision-making problems under uncertainty, with broad applications from preventive maintenance to event-triggered control can be modeled in the framework of semi-Markov decision processes (SMDPs). Unlike Markov decision processes (MDPs), SMDPs are underexplored in the online and reinforcement learning (RL) settings. In this paper, we extend the well-known deterministic policy gradient (DPG) theorem in MDPs to SMDPs under average-reward criterion. The existing stochastic policy gradient methods not only require, in general, a large number of samples for training, but they also suffer from high variance in the gradient estimation when applied to problems with deterministic optimal policy. Our DPG method can potentially remedy these issues. On the basis of this method and depending on the choice of a critic, different actor–critic algorithms can easily be developed in the RL setup. We present two example actor–critic algorithms. Both algorithms employ our developed policy gradient theorem for their actors, but use two different critics; one uses a simple SARSA update while the other one uses the same on-policy update but with compatible function approximators. We demonstrate the efficacy of our method both mathematically and via simulations.

Similar Papers
  • Conference Article
  • Citations26

MPC-based Reinforcement Learning for a Simplified Freight Mission of Autonomous Surface Vehicles

  • Dec 14, 2021
  • Wenqi Cai +4
  • Research Article
  • Citations6

A unified approach to time-aggregated Markov decision processes

  • Feb 04, 2016
  • Automatica
  • Yanjie Li +1
  • Research Article
  • Citations2

VMDP-program for solving optimality problems in vector criterion Markov and semi-Markov decision processes

  • Jan 01, 1991
  • Optimization
  • J Novák
  • Research Article
  • Citations2

A Multi-Step Reinforcement Learning Algorithm

  • Dec 06, 2010
  • Applied Mechanics and Materials
  • Zhi Cong Zhang +4
  • Conference Article
  • Citations4

On Robust and Adaptive Fidelity Selection for Human-in-the-loop Queues

  • Jun 29, 2021
  • Piyush Gupta +1
  • Book Chapter
  • Citations6

Soft Actor-Critic-Based DAG Tasks Offloading in Multi-access Edge Computing with Inter-user Cooperation

  • Jan 01, 2022
  • Pengbo Liu +4
  • Research Article
  • Citations22

A deep recurrent Q network towards self‐adapting distributed microservice architecture

  • Nov 28, 2019
  • Software: Practice and Experience
  • Basel Magableh +1
  • PDF
  • Research Article
  • Citations17

A novel approach to locomotion learning: Actor-Critic architecture using central pattern generators and dynamic motor primitives.

  • Oct 02, 2014
  • Frontiers in Neurorobotics
  • Cai Li +2
  • Conference Article
  • Citations2

Multi-Objective Policy Gradients with Topological Constraints

  • Oct 23, 2022
  • Kyle Hollins Wray +3
  • Research Article

The integration path of new generation information technology and ideological and political education in colleges and universities

  • Jan 01, 2024
  • Applied Mathematics and Nonlinear Sciences
  • Hui Tong +1
  • Research Article
  • Citations3

A constrained optimization perspective on actor–critic algorithms and application to network routing

  • May 05, 2016
  • Systems & Control Letters
  • Prashanth L.A +3
  • Research Article
  • Citations10

Time-average optimal constrained semi-Markov decision processes

  • Jun 01, 1986
  • Advances in Applied Probability
  • Frederick J Beutler +1
  • PDF
  • Research Article
  • Citations28

Learning to maximize reward rate: a model based on semi-Markov decision processes

  • May 23, 2014
  • Frontiers in Neuroscience
  • Arash Khodadadi +2
  • Research Article
  • Citations2

Метод синтеза нейронных регуляторов для линейных объектов

  • Dec 18, 2020
  • Science Bulletin of the Novosibirsk State Technical University
  • Dmitry Romannikov
  • Conference Article
  • Citations5

Optimal Management of Rechargeable Biosensors in Temperature-Sensitive Environments

  • Sep 01, 2010
  • Yahya Osais +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.