Autonomous agents are occupying more roles in our world than ever. They are present as AI in games, decide on which ads users see on the internet, and are even considered in more impactful environments such as finances, health, and perhaps traffic. As our expectations grow, their responsibilities become increasingly complex. To meet these demands, agents must reason over the current and potential future states of the environment. Reasoning over the state is not limited to understanding the latest observation but extends to inference over the hidden part of the environment. A natural solution is to maintain a probability distribution, or belief, over the state. Similarly, regarding future states, intelligent behavior considers the likelihood of multiple outcomes far into the future, also called planning. Both aspects of decision-making require full knowledge of the behavior, or dynamics, of the environment. These dynamics are used in planning to simulate future interactions. Similarly, imagined simulations of the system allow for reasoning about its current state. For example, a self-driving car may want to reason about the behavior of pedestrians for its decision-making (or indeed to estimate the likelihood of one being hidden from view!). Typically the dynamics are not fully known and, thus, these techniques are not directly applicable in practice. When the dynamics are not available, then the best thing we can do is to provide a prior over the unknown quantities instead. For example, even if there is no fully accurate model of pedestrians, we may provide a distribution over possible behaviors. This approach, called Bayesian reinforcement learning, allows us to provide the agent with expert knowledge. With this setup, it is now possible to update our understanding of the unknown quantities of the system when new data becomes available. The result is an inference problem where we maintain a belief over both the state and the dynamics of the environment. This Bayesian reinforcement learning approach was best formalized in the literature with the Bayes-adaptive models. Unfortunately, although elegant in theory, the approach had some limitations and saw little adoption in practice. In particular, no planner had been developed to allow for efficient action selection. Second, the type of prior knowledge that the proposed models were able to capture restricted their usage to small problems. Variants with more complex representations required strong priors and the structure of the dynamics would be assumed known. This thesis advances Bayesian partially observable reinforcement learning to non-trivial domains with several contributions. In particular, it discusses a holistic definition of the Bayesian inference problem, improved planning algorithms, and scalable model representations plus approaches for their approximation. First, we visualize the inference problem over state and dynamics with a graphical model and exploit its structure to define a novel posterior derivation. Not only does this derivation unify previous work into a single recipe, but also opens the door to other (machine learning) approaches. The resulting framework is formalized as the general Bayes-adaptive Markov decision process'' (GBA-POMDP). The second contribution is a family of efficient planning algorithms that were the first technique that made it possible to do decision-making in the Bayes-adaptive models. These planners are specializations of Monte-Carlo tree search, which are sample-based methods that approximate the value of future actions through simulations. In particular, we combine additional sampling approximations with inference simplifications to make reasoning over the complex GBA-POMDP possible. Third, we discuss two sophisticated models that move away from the tabular representations that were assumed in previous work. We show how a Bayes-net representation allows us to model and learn structure in the dynamics of the system. This includes the usage of an intricate and targeted sampling scheme to avoid the collapse of the posterior approximation. Lastly, we show the practical use of the GBA-POMDP withBayes-Adaptive Deep Dropout Reinforcement learning (BADDr)'', which employs neural networks to model the dynamics of the environment. BADDr, with the expressiveness and scalability of neural networks, showcases how Bayesian reinforcement learning can be both principled and practical.--Author's abstract
Read more