- Research Article
11
- 10.1016/s0142-0615(97)00066-5
Optimal load clipping with time of use rates
- May 01, 1998
- International Journal of Electrical Power and Energy Systems
- Isto Aho + 3 more +3
Optimal load clipping with time of use rates
In this paper, we present a modified dynamic programming (DP) method. The method is basically the same as the value iteration method (VI), a representative DP method, except the preprocess of a system's state transition model for reducing its complexity, and is called the dynamic programming on reduced models (DPRM). That reduction is achieved by imaginarily considering causes of the probabilistic behavior of a system, and then cutting off some causes with low occurring probabilities. In computational illustrations, VI, DPRM, and the real-time Q-learning method (RTQ) are applied to elevator operation problems, which can be modeled by using Markov decision processes. The results show that DPRM can compute quasi-optimal value functions which bring more effective allocations of elevators than value functions by RTQ in less computational times than VI. This characteristic is notable when the traffic pattern is complicated.
Optimal load clipping with time of use rates
Optimal load clipping with time of use rates
Approximate Dynamic Programming
In any complex or large scale sequential decision making problem, there is a crucial need to use function approximation to represent the relevant functions such as the value function or the policy. The Dynamic Programming (DP) and Reinforcement Learning (RL) methods introduced in previous chapters make the implicit assumption that the value function can be perfectly represented (i.e. kept in memory), for example by using a look-up table (with a finite number of entries) assigning a value to all possible states (assumed to be finite) of the system. Those methods are called exact because they provide an exact computation of the optimal solution of the considered problem (or at least, enable the computations to converge to this optimal solution). However, such methods often apply to toy problems only, since in most interesting applications, the number of possible states is so large (and possibly infinite if we consider continuous spaces) that a perfect representation of the function at all states is impossible. It becomes necessary to approximate the function by using a moderate number of coefficients (which can be stored in a computer), and therefore extend the range of DP and RL to methods using such approximate representations. These approximate methods combine DP and RL methods with function approximation tools.
Read moreSolving Markov Decision Processes via Simulation
This chapter presents an overview of simulation-based techniques useful for solving Markov decision processes (MDPs). MDPs model problems of sequential decision-making under uncertainty, in which decisions made in each state collectively affect the trajectory of the states visited by the system over a time horizon of interest. Traditionally, MDPs have been solved via dynamic programming (DP), which requires the transition probability model that is difficult to derive in many realistic settings. The use of simulation for solving MDPs allows us to bypass the transition probability model and solve large-scale MDPs considered intractable to solve by traditional DP. The simulation-based methodology for solving MDPs, which like DP is also rooted in the Bellman equations, goes by names such as reinforcement learning, neuro-DP, and approximate or adaptive DP. We begin with a description of algorithms for infinite-horizon discounted reward MDPs, followed by the same for infinite-horizon average reward MDPs. Then we present a discussion on finite-horizon MDPs. For each problem considered, we present a step-by-step description of a selected group of algorithms. In making this selection, we have attempted to blend the old and the classical with more recent developments. Finally, after touching upon extensions and convergence theory, we conclude with a brief summary of some applications and directions for future research.
Read moreModeling of Marketing Processes Using Markov Decision Process Approach
In this paper, an application of Markov Decision Processes (MDP) for modeling selected marketing process is presented. The process is converted into MDP model, where states of the MDP are determined by a configuration of state vector. Elements of the state vector represent most important attributes of the customer in the modeled process. Movement between the states is determined by actions of the customer. In constructed MDP model, individual states with assigned initial reward values then represent consequences of action chosen by the customer resulting in either incresing or reducing the revenue following moving into these states. Value iteration method is then used for computation of the expected final rewards for each state. Based on provided realistic data, customer behavior is analyzed and the best course of action is proposed. Model suitability for future predictions of desired action outcome rate is discussed as well.
Read moreSimplifying Dynamic Programming via Tabling
In the dynamic programming paradigm the value of an optimal solution is recursively defined in terms of optimal solutions to subproblems. Such dynamic programming definitions can be very tricky and error-prone to specify. This paper presents a novel, elegant method based on tabled logic programming that simplifies the specification of such dynamic programming solutions. Our method introduces a new mode declaration for tabled predicates. The arguments of each tabled predicate are divided into indexed and non-indexed ones so that tabled predicates can be regarded as functions: indexed arguments represent input values and non-indexed arguments represent output values. The non-indexed arguments in a tabled predicate can be further declared to be aggregate, e.g., the minimum, so that while generating answers, the global table will dynamically maintain the smallest value for that argument. This mode declaration scheme, coupled with recursion, provides a considerably easy-to-use method for dynamic programming: there is no need to define the value of an optimal solution recursively, instead, defining a general solution suffices. The optimal value as well as its corresponding concrete solution can be derived implicitly and automatically using tabled logic programming systems. Experimental results are shown to indicate that the mode declaration improves both time and space performances in solving dynamic programming problems on tabled LP systems. Additionally, our mode declaration scheme provides an alternative implementation vehicle for preference logic programming.
Read moreSimplifying dynamic programming via mode‐directed tabling
In the dynamic programming paradigm the value of an optimal solution is recursively defined in terms of optimal solutions to subproblems. Such dynamic programming definitions can be tricky and error‐prone to specify. This paper presents an elegant method based on tabled logic programming (TLP) that simplifies the specification of such dynamic programming solutions. Our method introduces a new mode declaration for tabled predicates. The arguments of each tabled predicate are divided into indexed and non‐indexed arguments so that tabled predicates can be regarded as functions: indexed arguments represent input values and non‐indexed arguments represent output values. The non‐indexed arguments in a tabled predicate can be further declared to be aggregated, for example, the minimum, so that while generating answers, the global table will dynamically maintain the smallest value for that argument. This mode‐declaration scheme, coupled with recursion, provides an easy‐to‐use method for dynamic programming: there is no need to define the value of an optimal solution recursively, as the definition of a general solution suffices. The optimal value as well as its corresponding concrete solution can be derived implicitly and automatically using tabled logic programming systems. Our experimental results show that mode declarations improve performance in solving dynamic programming problems on TLP systems. Copyright © 2007 John Wiley & Sons, Ltd.
Read moreStudy on Water Quantity Allocation Optimization for Single Main Canal in Large-Scale Irrigation Area Based on DP Method
The mathematical model of optimal water quantity allocation for a single main canal in a large-scale irrigation area was constructed that took the minimal sum of the squared deviation of water shortage for water receiving areas controlled by the single main canal in one given irrigation period as the study target, and the total irrigation quantity of the single main canal as a constraint condition. Taking the optimal allocation of water quantity of each branch canal as decision variables, and several branch canals under the irrigation sequence of the main canal as a state variable, this model was solved by the one-dimensional dynamic programming (DP) method, by which the minimal water shortage and corresponding optimal water quantity allocation of each branch canal was calculated. The proposed method could provide a decision-making reference for optimal water resources allocation of single main canal irrigation areas, and also provide the theoretical basis for optimal water quantity allocation of a main canal with rotation irrigation by strips or with segmented rotation irrigation mode in China’s large-scale irrigation areas. Taking Hengliu Main Canal of Zhouqiao Irrigation Area in Jiangsu Province as a study case, optimization results showed that in a medium drought year (p = 75%) and a special drought year (p = 95%), minimal water shortage for water receiving areas controlled by Hengliu Main Canal was respectively 2.57 × 104 m3 and 23.31 × 104 m3 during the ponding period of rice. The corresponding water quantity allocation for each branch canal has reflected a compellent model solution precision and efficiency.
Read moreHorizontal combinations of online and offline approximate dynamic programming for stochastic dynamic vehicle routing
Stochastic and dynamic vehicle routing problems gain increasing attention in the research community. In these problems, routing plans are dynamically updated based on realizations of stochastic information. Due to the complexity of the corresponding Markov decision processes (MDPs), the calculation of optimal policies for these problems is usually not possible and researchers draw on heuristical methods of approximate dynamic programming (ADP). These methods use simulation to approximate the value of a state and decision in the MDP. The simulations are either conducted offline or online. Offline methods such as value function approximations (VFAs) generally neglect the full detail of the state space due to aggregation. Online methods such as rollout algorithms (RAs) are often not able to capture decision and transition space sufficiently due to runtime limitations. In this paper, we alleviate this tradeoff by combining two methods of ADP, an online RA and an offline VFA in two ways. In addition to the integration of the VFA as a base policy into the online RA to strengthen the RA’s simulations, we also limit the RA’s simulation horizon, estimating the remaining reward-to-go again via the VFA. For two stochastic dynamic routing problems from the literature, we show how this combination outperforms state-of-the-art solutions while simultaneously reducing the required time for online calculations.
Read moreMultidimensional Dynamic Programming for Homology Search on Distributed Systems
Alignment problems in computational biology have been focused recently because of the rapid growth of sequence databases. By computing alignment, we can understand similarity among the sequences. Dynamic programming is a technique to find optimal alignment, but it requires very long computation time. We have shown that dynamic programming for more than two sequences can be efficiently processed on a compact system which consists of an off-the-shelf FPGA board and its host computer (node). The performance is, however, not enough for comparing long sequences. In this paper, we describe a computation method for the multidimensional dynamic programming on distributed systems. The method is now being tested using two nodes connected by Ethernet. According to our experiments, it is possible to achieve 5.1 times speedup with 16 nodes, and more speedup can be expected for comparing longer sequences using more number of nodes. The performance is affected only a little by the data transfer delay when comparing long sequences. Therefore, our method can be mapped on any kinds of networks with large delays.
Read moreDynamic Programming Algorithms in Global Optimization
Dynamic Programming (DP) is a useful approach to multi-stage decision problems. On the basis of Bellman’s optimality principle such a problem can be decomposed into subproblems through recursive formulae. Bellman (1957), 1957a and Dantzig (1957) introduced a DP method for integer-variable linear programs. This method yields pseudo-polynomial algorithms when the number of constraints is fixed (Papadimitriou (1981)). In the class of linear programs, problems with a staircase structure can be efficiently treated by DP methods which can be interpreted as iterative predictor-corrector processes using a pricing mechanism for subproblems at individual stages (e.g., Dantzig (1963), Ho and Manne (1974), Ho (1978), Ho and Loute (1980), Abrahamson (1981) . Various applications of DP have also been developed in production planning (Manne (1958) , Wagner and Whitin (1959), Clark and Scarf (1960), Veinott (1965), 1969), Bessler and Veinott (1966), Zangwill (1965), 1966) (1969), Konno (1973), (1988), Bitran and Yanasse (1982), Bitran et al. (1984),... In particular, polynomial algorithms have been obtained for concave cost lotsizing problems and their extensions (e.g., Wagner and Whitin (1959), Zangwill (1965), (1966), (1959), Dreyfus and Law (1977).
Read moreDecision Analysis Based on Enterprise Production Process
This paper mainly studies the decision-making problems encountered by enterprises in the production process, including whether to inspect parts and finished products, and how to handle unqualified parts and finished products. By applying operations research and statistical knowledge, and adopting dynamic programming and one-sided hypothesis testing methods, relevant decision-making problems are effectively solved. In this paper, the authors propose a sampling inspection method to parameterize the defect rate of parts and optimize decision-making schemes for different production stages. Finally, the study analyzes the impact of various decisions on the economic benefits of enterprises, constructs corresponding models using dynamic programming, and derives optimal solutions under various decisions. For the first problem, based on the basic principles of statistics, this paper uses one-sided hypothesis testing to detect whether the defect rate of parts exceeds the nominal value. In the case of a large sample size, the central limit theorem is cleverly applied to approximate the binomial distribution to the normal distribution, thereby simplifying the calculation process. And based on the principle of sampling inspection, detailed judgments are made on whether to accept or reject parts under different confidence levels. When studying the second problem, this paper cleverly adopts the dynamic programming method to conduct detailed decision analysis on the three key stages of the enterprise production process: part inspection, finished product inspection, and unqualified product handling. And through the reverse analysis method, starting from the final product, the decision-making of each stage is gradually optimized to ensure that the risk is minimized while controlling costs. By evaluating the relationship between inspection costs and potential losses, as well as the specific impact of different handling methods for unqualified products on the economic benefits of enterprises, the goal is to maximize economic benefits while ensuring product quality. The third problem further enhances the complexity of decision-making based on the second problem, considering multiple processes and multiple parts to optimize decision-making in the multi-stage production process. This problem still adopts the reverse analysis method of dynamic programming to construct a more complex dynamic programming model, comprehensively considering the inspection costs, unqualified product handling costs, and potential market risks of each production stage. In the process of model construction, in-depth analysis is conducted on how to effectively handle parts, semi-finished products, unqualified semi-finished products, finished products, and unqualified finished products at different production stages, including different decisions for each stage. Through careful analysis, it provides enterprises with an optimal decision-making scheme for multiple processes and multiple parts.
Read moreThe application of dynamic programming to slope stability analysis
The applicability of the dynamic programming method to two-dimensional slope stability analyses is studied. The critical slip surface is defined as the slip surface that yields the minimum value of an optimal function. The only assumption regarding the shape of the critical slip surface is that the surface is an assemblage of linear segments. Stresses acting along the critical slip surface are computed using a finite element stress analysis. Assumptions associated with limit equilibrium methods of slices related to the shape of the critical slip surface and the relationship between interslice forces are no longer required. A computer program named DYNPROG was developed based on the proposed analytical procedure, and numerous example problems have been analyzed. Results obtained when using DYNPROG were compared with those obtained when using several well-known limit equilibrium methods. The comparisons demonstrate that the dynamic programming method provides a superior solution when compared with conventional limit equilibrium methods. Analyses conducted also show that factors of safety computed when using the dynamic programming method are generally slightly lower than those computed using conventional limit equilibrium methods of slices; however, as Poisson's ratio approaches 0.5, the computed factors of safety from the dynamic programming method and the limit equilibrium method appear to become similar.Key words: dynamic programming, slope stability, stress analysis, optimization theory, limit equilibrium methods of slices.
Read moreBounding procedure for stochastic dynamic programs with application to the perimeter patrol problem
One often encounters the curse of dimensionality in the application of dynamic programming to determine optimal policies for controlled Markov chains. In this paper, we provide a method to construct sub-optimal policies along with a bound for the deviation of such a policy from the optimum via a linear programming approach. The state-space is partitioned and the optimal cost-to-go or value function is approximated by a constant over each partition. By minimizing a positive cost function defined on the partitions, one can construct an approximate value function which also happens to be an upper bound for the optimal value function of the original Markov Decision Process (MDP). As a key result, we show that this approximate value function is independent of the positive cost function (or state dependent weights; as it is referred to, in the literature) and moreover, this is the least upper bound that one can obtain; once the partitions are specified. We apply the linear programming approach to a perimeter surveillance stochastic optimal control problem; whose structure enables efficient computation of the upper bound.
Read moreCharging Scheduling of Enterprise Private Parking-lot with Renewable Power: A simulation method based on Markov Decision Processes
With the development of renewable energy and electric vehicles (EVs), using renewable energy charging for EVs becomes an effective way to reduce environment pollution. In this paper, we study a charging scheduling problem in the enterprise private parking-lot considering the uncertainties of renewable energy and charging demand of EVs. In this model, the stochastic multistage decision problem is described as Markov decision processes (MDPs). In addition, a simulation-based dynamic programming (SBDP) method is proposed to get the optimal strategy and expected cost of purchasing electricity from the power grid. Finally, a real case study is analyzed. The results show that the method is applicable to various scale of EVs well and the cost of electricity purchasing has been reduced in the range of 9% and 47% depending on the scale of charging piles.
Read moreOnline adaptive dynamic programming algorithm for linear fixed-time impulse hybrid systems
This paper studies an online adaptive dynamic programming (ADP) algorithm for linear fixed-time impulse systems. The ADP algorithm follows the value iteration method, updating the value approximation and the control law iteratively. In the algorithm, a single network structure is adopted. The single critic network approximates the objective value function. Online training of the value function and the control laws are implemented at each iteration. The network approximated value function converges to the optimal one of the impulse system. Simulations are also presented for the validity of the presented algorithm.
Read more