• Home
  • Search
  • An Efficient Simulation-Based Policy Improvement with Optimal Computing Budget Allocation Based on Accumulated Samples
  • Open Access IconOpen Access
  • Cite Icon1
  • https://doi.org/10.3390/electronics11071141Copy DOI Icon

An Efficient Simulation-Based Policy Improvement with Optimal Computing Budget Allocation Based on Accumulated Samples

Show More
  • Abstract
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Markov decision processes (MDPs) are widely used to model stochastic systems to deduce optimal decision-making policies. As the transition probabilities are usually unknown in MDPs, simulation-based policy improvement (SBPI) using a base policy to derive optimal policies when the state transition probabilities are unknown is suggested. However, estimating the Q-value of each action to determine the best action in each state requires many simulations, which results in efficiency problems for SBPI. In this study, we propose a method to improve the overall efficiency of SBPI using optimal computing budget allocation (OCBA) based on accumulated samples. Previous works have mainly focused on improving SBPI efficiency for a single state and without using the previous simulation samples. In contrast, the proposed method improves the overall efficiency until an optimal policy can be found in consideration of the state traversal property of the SBPI. The proposed method accumulates simulation samples across states to estimate the unknown transition probabilities. These probabilities are then used to estimate the mean and variance of the Q-value for each action, which allows the OCBA to allocate the simulation budget efficiently to find the best action in each state. As the SBPI traverses the state, the accumulated samples allow appropriate allocation of OCBA; thus, the optimal policy can be obtained with a lower budget. The experimental results demonstrate the improved efficiency of the proposed method compared to previous works.

Loading PDF

Similar Papers
  • Book Chapter
  • Citations3

A Rollout Algorithm for Multichain Markov Decision Processes with Average Cost

  • Jan 01, 2009
  • Tao Sun +2
  • Research Article
  • Citations100

Minimizing Risk Models in Markov Decision Processes with Policies Depending on Target Values

  • Mar 01, 1999
  • Journal of Mathematical Analysis and Applications
  • Congbin Wu +1
  • Research Article
  • Citations33

Using simulation and optimisation to characterise durations of emergency department service times with incomplete data

  • Jul 14, 2016
  • International Journal of Production Research
  • Hainan Guo +4
  • Conference Article
  • Citations4

Solving Stationary and Stochastic Point Location Problem with Optimal Computing Budget Allocation

  • Oct 01, 2015
  • Junqi Zhang +2
  • Conference Article
  • Citations1

A hybrid search algorithm with optimal computing budget allocation for resource allocation problem

  • Dec 08, 2013
  • James T Lin +1
  • Conference Article
  • Citations22

Resampling in Particle Swarm Optimization

  • Jun 01, 2013
  • Juan Rada-Vilela +2
  • Conference Article
  • Citations3

Cost Minimization for Admission Control in Bandwidth Asymmetry Wireless Networks

  • Jun 01, 2007
  • X Yang +1
  • Conference Article
  • Citations7

Optimal Computing Budget Allocation Under Correlated Sampling

  • Apr 05, 2005
  • M.C Fu +3
  • PDF
  • Research Article
  • Citations7

An Effective Adjustment to the Integration of Optimal Computing Budget Allocation for Particle Swarm Optimization in Stochastic Environments

  • Jan 01, 2020
  • IEEE Access
  • Seon Han Choi +1
  • Research Article
  • Citations134

Managing Wind‐Based Electricity Generation in the Presence of Storage and Transmission Capacity

  • Apr 01, 2019
  • Production and Operations Management
  • Yangfang (Helen) Zhou +3
  • PDF
  • Research Article
  • Citations7

Integrating Multiple Policies for Person-Following Robot Training Using Deep Reinforcement Learning

  • Jan 01, 2021
  • IEEE Access
  • Chandra Kusuma Dewa +1
  • Research Article
  • Citations18

Stochastic optimal control of agrochemical pollutant loads in reservoirs for irrigation

  • May 27, 2016
  • Journal of Cleaner Production
  • Goden Mabaya +2
  • Research Article
  • Citations5

Energy-Efficient Joint Pushing and Caching Based on Markov Decision Process

  • Jun 01, 2019
  • IEEE Transactions on Green Communications and Networking
  • Hoshyar Mohammed +2
  • PDF
  • Research Article
  • Citations1

Three-layer model for the control of epidemic infection over multiple social networks

  • Apr 29, 2023
  • Sn Applied Sciences
  • Ali Nasir
  • Research Article
  • Citations11

Dynamic selling of quality-graded products under demand uncertainties

  • Mar 15, 2011
  • Computers & Industrial Engineering
  • Hua-Hsuan Wu +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.