- Research Article
- 10.1287/opre.1110.0925
Contributors
- Feb 01, 2011
- Operations Research
- Sandro Bosio
Contributors
Approximate dynamic programming based approach to process control and scheduling
Contributors
Contributors
Approximate dynamic programming: Application to process supply chain management
Many continuous process industries are required to operate multiproduct supply chains under significant uncertainty in market demands and prices, and mathematical programming formulations have been developed for the deterministic version of the problem. However, available extensions to the stochastic case cannot be readily used to obtain an optimal operating policy in practice because the uncertainty results in an explosion of the scenarios to be considered over an extended time horizon. In this article we represent the uncertainty through Markov chains and use an approach based on stochastic dynamic programming (DP), which can generate a dynamic operating policy that incorporates information about the uncertainty in the problem at each time step. DP is computationally infeasible because of the size of the state and action space for realistic problems. We propose a novel stochastic optimization algorithmic framework called “DP in a heuristically restricted state space.” The confined state space is generated by simulating various potential scenarios under centralized dynamic inventory and production policy generated by combining static local inventory policy heuristics. The resulting DP policy responds to the time‐varying demand for products by appropriately stitching together decisions made by the local static policies. This is a new formulation for effectively combining heuristic policies. © 2006 American Institute of Chemical Engineers AIChE J, 2006
Read moreSolving Markov Decision Processes via Simulation
This chapter presents an overview of simulation-based techniques useful for solving Markov decision processes (MDPs). MDPs model problems of sequential decision-making under uncertainty, in which decisions made in each state collectively affect the trajectory of the states visited by the system over a time horizon of interest. Traditionally, MDPs have been solved via dynamic programming (DP), which requires the transition probability model that is difficult to derive in many realistic settings. The use of simulation for solving MDPs allows us to bypass the transition probability model and solve large-scale MDPs considered intractable to solve by traditional DP. The simulation-based methodology for solving MDPs, which like DP is also rooted in the Bellman equations, goes by names such as reinforcement learning, neuro-DP, and approximate or adaptive DP. We begin with a description of algorithms for infinite-horizon discounted reward MDPs, followed by the same for infinite-horizon average reward MDPs. Then we present a discussion on finite-horizon MDPs. For each problem considered, we present a step-by-step description of a selected group of algorithms. In making this selection, we have attempted to blend the old and the classical with more recent developments. Finally, after touching upon extensions and convergence theory, we conclude with a brief summary of some applications and directions for future research.
Read moreAn S-graph Based Approach for Multi-Mode Resource-Constrained Project Scheduling with Time-Varying Resource Capacities
The Resource-Constrained Project Scheduling Problem (RCPSP) is a general problem class of scheduling problems. It even contains the well-known job shop scheduling problem as a special case. In these problems, jobs demand different amounts from multiple resources at once. The goal is to minimize completion time while satisfying resource capacity constraints throughout the schedule. The RCPSP has several variations and generalizations, one of which is the multi-mode problem. Multiple operation modes are given for jobs, differing in resource usages and execution times. Another generalization is where resources have time-varying capacities instead of a constant limit. The S-graph framework was developed for batch process scheduling and used successfully in various applications. The advantages of this approach motivate research to extend its capabilities to more general problem classes and apply it in various specialized case studies. The authors have previously presented an extension of the S-graph framework for RCPSP and its multi-mode generalization. In this work, a further extension is proposed for handling time-varying capacities. The method relies on a model transformation technique and recent algorithmic developments of the framework.
Read moreApplying unweighted least‐squares based techniques to stochastic dynamic programming: theory and application
Big data and the curse of dimensionality are common vocabularies that researchers in different communities have recently been dealing with, e.g. dynamic programming (DP) in automatic control system society. A novel unweighted sampled based least square projection approach is proposed in this study to address the issue of the large state space in the DP optimisation problem. The method, in particular, takes into account both contraction mapping and monotonicity properties of the DP algorithm for value function approximation. Specifically, the batch of samples are gathered by uniform probability distribution at first, and an unweighted LS sub‐problem in the subspace is solved. As the case study, a new Markov decision process model associated with a resource allocation problem is considered to illustrate the technique and evaluate its effectiveness. It is noted that the approach can be employed for different applications as well. Moreover, a MATLAB based software is developed to implement and examine different parts of the proposed method. Simulation examples are considered to support the results of the approach via developed software. The idea makes a connection between the recent advances in big data analysis and approximate DP as well.
Read moreA Least-Squares Temporal Difference based method for solving resource allocation problems
A Least-Squares Temporal Difference based method for solving resource allocation problems
Approximate dynamic programming for an energy-efficient parallel machine scheduling problem
Approximate dynamic programming for an energy-efficient parallel machine scheduling problem
The Sample Analysis Machine Scheduling Problem: Definition and comparison of exact solving approaches
The Sample Analysis Machine Scheduling Problem: Definition and comparison of exact solving approaches
Rejoinder—The Languages of Stochastic Optimization
The theme of my feature article is the crossfertilization of ideas from artificial intelligence and operations research, although it is important to emphasize that some of the key ideas of approximate dynamic programming also draw heavily on developments whose origins should be attributed to control theory. I am very grateful for the insightful remarks of John N. Tsitsiklis, well known for his many contributions to approximate dynamic programming and his coauthorship of a seminal book (Bertsekas and Tsitsiklis 1996), and Andrzej Ruszczynski, one of the very top names in the field of stochastic programming who, in addition to numerous research papers, coedited the volume on stochastic programming in the prestigious Handbook on Operations Research and Management Science series (Ruszczynski and Shapiro 2003). I also benefited from extensive comments by Andy Barto who, with Richard Sutton, coauthored the book Reinforcement Learning, which is the bible of this hugely popular field (Sutton and Barto 1998). Both sets of comments add a number of useful insights. I am going to focus my final remarks less on the specifics of their remarks and more on the differences in perspectives they bring from their home communities. I also need to note that there is no such thing as the “operations research perspective”— rather, there are at least four major subcommunities within operations research that deal directly with stochastic optimization: Markov decision processes (MDPs) (as represented by Puterman 1994), stochastic programming (Kall and Wallace 1994, Higle and Sen 1996, and Birge and Louveaux 1997 are major references), decision theory (primarily decision trees, as in Skinner 1999), and simulation-optimization (see, for example, Swisher et al. 2000, Fu 2002). Tsitsiklis (2010) notes the strong similarities between operations research and the field of systems and control theory. In drawing on this conclusion, he is focusing on the central paradigm of optimizing a known objective function with known system dynamics. Artificial intelligence (AI), by contrast, is often addressing the problem of learning how people behave (or more fundamentally, how the brain solves problems), often in the context of applications where we may not understand why someone is making a particular decision (hence, we may not know the reward function). Also, these applications often take place in an online setting where we are observing events from an external, physical system. As a result, we may be able to observe the next state we land in, but we do not necessarily know the physics of why we landed there. The first difference someone will notice between these communities appears purely cosmetic. The AI and MDP communities use state s and action a; elsewhere, operations research will typically use x for a decision, expecting it to be a high-dimensional vector. In control theory, the state is x and the decision (control) is u. Behind these cosmetic differences are problem classes with associated empirical properties and computational expectations. A 10-dimensional control problem can be much harder than an OR problem with 10,000 dimensions. A second and significant difference between the communities is how they represent the dynamics of the system. In the MDP community, it is absolutely routine to start by knowing the transition matrix p s′ s x , which gives the probability of landing in (typically discrete) state s′, given that we are in state s and take action x. However, there are many problems where this matrix is either simply unknown or computationally intractable (for example, if the number of states is too large). Computing the transition matrix requires first knowing the transition function, which might be written as (using my notation)
Read moreA novel ADP based model-free predictive control
Dynamic programming is a very useful tool in solving optimization and optimal control problems. Here, the Approximate Dynamic Programming (ADP) and the notion of neural networks based predictive control are combined with a model-free control method based on SPSA (Simultaneous perturbation stochastic approximation), and a novel ADP based model-free predictive control strategy for nonlinear systems is proposed. Dynamic programming is used to adjust the control parameters in the novel model-free control method and the notion of predictive control is introduced to modify the whole control structure. Finally, the proposed ADP based model-free predictive control strategy is applied to solve nonlinear tracking problems and the effectiveness of this novel control method is fully illustrated though simulation tests on two typical nonlinear systems.
Read morePerformance Evaluation of Direct Heuristic Dynamic Programming using Control-Theoretic Measures
Approximate dynamic programming (ADP) has been widely studied from several important perspectives: algorithm development, learning efficiency measured by success or failure statistics, convergence rate, and learning error bounds. Given that many learning benchmarks used in ADP or reinforcement learning studies are control problems, it is important and necessary to examine the learning controllers from a control-theoretic perspective. This paper makes use of direct heuristic dynamic programming (direct HDP) and three typical benchmark examples to introduce a unique analytical framework that can be applied to other learning control paradigms and other complex control problems. The sensitivity analysis and the linear quadratic regulator (LQR) design are used in the paper for two purposes: to quantify direct HDP performances and to provide guidance toward designing better learning controllers. The use of LQR however does not limit the direct HDP to be a learning controller that addresses nonlinear dynamic system control issues. Toward this end, applications of the direct HDP for nonlinear control problems beyond sensitivity analysis and the confines of LQR have been developed and compared whenever appropriate to an LQR design.
Read moreApproximate Dynamic Programming
In any complex or large scale sequential decision making problem, there is a crucial need to use function approximation to represent the relevant functions such as the value function or the policy. The Dynamic Programming (DP) and Reinforcement Learning (RL) methods introduced in previous chapters make the implicit assumption that the value function can be perfectly represented (i.e. kept in memory), for example by using a look-up table (with a finite number of entries) assigning a value to all possible states (assumed to be finite) of the system. Those methods are called exact because they provide an exact computation of the optimal solution of the considered problem (or at least, enable the computations to converge to this optimal solution). However, such methods often apply to toy problems only, since in most interesting applications, the number of possible states is so large (and possibly infinite if we consider continuous spaces) that a perfect representation of the function at all states is impossible. It becomes necessary to approximate the function by using a moderate number of coefficients (which can be stored in a computer), and therefore extend the range of DP and RL to methods using such approximate representations. These approximate methods combine DP and RL methods with function approximation tools.
Read moreSimulation-Based Algorithms for Markov Decision Processes: Monte Carlo Tree Search from AlphaGo to AlphaZero
AlphaGo and its successors AlphaGo Zero and AlphaZero made international headlines with their incredible successes in game playing, which have been touted as further evidence of the immense potential of artificial intelligence, and in particular, machine learning. AlphaGo defeated the reigning human world champion Go player Lee Sedol 4 games to 1, in March 2016 in Seoul, Korea, an achievement that surpassed previous computer game-playing program milestones by IBM’s Deep Blue in chess and by IBM’s Watson in the U.S. TV game show Jeopardy. AlphaGo then followed this up by defeating the world’s number one Go player Ke Jie 3-0 at the Future of Go Summit in Wuzhen, China in May 2017. Then, in December 2017, AlphaZero stunned the chess world by dominating the top computer chess program Stockfish (which has a far higher rating than any human) in a 100-game match by winning 28 games and losing none (72 draws) after training from scratch for just four hours! The deep neural networks of AlphaGo, AlphaZero, and all their incarnations are trained using a technique called Monte Carlo tree search (MCTS), whose roots can be traced back to an adaptive multistage sampling (AMS) simulation-based algorithm for Markov decision processes (MDPs) published in Operations Research back in 2005 [Chang, HS, MC Fu, J Hu and SI Marcus (2005). An adaptive sampling algorithm for solving Markov decision processes. Operations Research, 53, 126–139.] (and introduced even earlier in 2002). After reviewing the history and background of AlphaGo through AlphaZero, the origins of MCTS are traced back to simulation-based algorithms for MDPs, and its role in training the neural networks that essentially carry out the value/policy function approximation used in approximate dynamic programming, reinforcement learning, and neuro-dynamic programming is discussed, including some recently proposed enhancements building on statistical ranking & selection research in the operations research simulation community.
Read moreComputationally efficient approximate dynamic programming for multi-site production capacity planning with uncertain demands
With globalization and rapid technological-economic development accelerating the market dynamics, consumers' demand is becoming more volatile and diverse. In this situation, capacity adjustment as an operational strategic decision plays a major role to ensure supply chain responsiveness while maintaining costs at a reasonable norm. This study contributes to the literature by developing computationally efficient approximate dynamic programming approaches for production capacity planning considering uncertainties and demand interdependence in a multi-factory multi-product supply chain setting. For this purpose, the k-Nearest-Neighbor-based Approximate Dynamic Programming and the Rolling-Horizon-based Approximate Dynamic Programming are developed to enable real-time decision support while ensuring the robustness of the outcomes in stochastic decision environments. Given the market volatilities in the Thin Film Transistor-Liquid Crystal Display industry, a real case from this sector is investigated to evaluate the applicability of the developed approach and provide insights for other industry situations. The developed method is less complex to implement, and numerical experiments showed that it is also computationally more efficient compared to Stochastic Dynamic Programming.
Read moreRobust Optimizers for Nonlinear Programming in Approximate Dynamic Programming
Many stochastic dynamic programming tasks in continuous action-spaces are tackled through discretization. We here avoid discretization; then, approximate dynamic programming (ADP) involves (i) many learning tasks, performed here by Support Vector Machines, for Bellman-function-regression (ii) many non-linear-optimization tasks for action-selection, for which we compare many algorithms. We include discretizations of the domain as well as other non-linear-programming tools in our experiments, so that by the way we compare optimization approaches and discretization methods. We conclude that robustness is strongly required in the non-linear optimizations in ADP, and experimental results show that (i) discretization is sometimes inefficient, but some specific discretization is very efficient for "bang-bang" problems (ii) simple evolutionary tools outperform quasi-random in a stable manner (iii) gradient-based techniques are much less stable (iv) for most high-dimensional “less unsmooth” problems Covariance-Matrix-Adaptation is first ranked.KeywordsGenetic AlgorithmSupport Vector MachineEvolutionary AlgorithmReinforcement LearningRobust OptimizerThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Read more