• Home
  • Search
  • Reinforcement Learning for Clinical Applications.
  • Open Access IconOpen Access
  • Cite Icon13
  • https://doi.org/10.2215/cjn.0000000000000084Copy DOI Icon

Reinforcement Learning for Clinical Applications.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Introduction Reinforcement learning formalizes the concept of learning from interactions.1 Broadly, reinforcement learning focuses on a setting in which an agent (decision maker) sequentially interacts with an environment that is partially unknown to them. At each stage, the agent takes an action and receives a reward. The objective of the agent is to maximize rewards accumulated in the long run. There are many situations in health care where decisions are made sequentially for which reinforcement learning approaches could prove useful for decision making. Throughout this article, we consider treatment prescription as an archetypical example to connect reinforcement learning concepts to a health care setting. In this setting, the care provider, the prescribed treatment, and the patients can be viewed as the agent, the action, and the environment, respectively, as depicted in Figure 1.Figure 1: Sequential treatment of AKI or CKD complications modeled as a reinforcement learning problem.Background In this section, with the objective of making reinforcement learning literature more accessible to a clinical audience, we briefly introduce related fundamental concepts and approaches. We refer the interested reader to Sutton and Barto1 for a comprehensive introduction to reinforcement learning. Markov Decision Processes Markov decision processes (MDPs) are a formalism of the sequential decision-making problem that has been central to the theoretical and practical advancements of reinforcement learning. In each stage of an MDP, the agent observes the state of the environment and takes an action, which, in turn, results in a change of the state. This change of state is assumed to be probabilistic with the next state being determined only by the preceding state, the chosen action, and the transition probability. The agent also receives a reward that is a function of the taken action, the preceding state, and the subsequent state. In an MDP, the objective of the agent is to maximize the return defined as the reward accumulated over a time horizon. In some applications, it is common to consider the horizon to be infinite, in which case the future rewards are discounted by a factor smaller than one. The selection of action by the agent on the basis of the observed state is known as the policy. More formally, a policy is a probabilistic mapping from states to each possible action. Because the policy and the reward are a function of the state, it is critical to estimate the utility of being in a certain state. More specifically, the value function is defined as the expected return starting from a given state under the chosen policy. Under this formalism, the objective of the agent is to find the optimal policy that maximizes the value function for all states. Reinforcement Learning Methods Action-value methods are a class of reinforcement learning methods in which the actions are chosen on the basis of the estimation of their long-term value. A prominent example of an action-value method is Q-learning in which the agent iteratively takes actions with the highest estimated values and updates the action-state value function on the basis of new observations. Policy gradient methods are another class of reinforcement learning methods that seek to optimize the policy directly instead of choosing actions on the basis of their respective estimated value. Such methods could be advantageous in health care applications that entail a large number of possible actions, e.g., when recommending a wide range of drug dosages or treatment options. Clinical Applications Reinforcement learning frameworks and methods are broadly applicable to clinical settings in which decisions are made sequentially. A prominent clinical application of reinforcement learning is for treatment recommendation, which has been studied across a variety of diseases and treatments including radiation and chemotherapy for cancer, brain stimulation for epilepsy, and treatment strategies for sepsis.2–5 In such treatment recommendation settings, a policy is commonly known as a dynamic treatment regime. There are various other clinical applications of reinforcement learning including diagnosis, medical imaging, and decision support tools (see refs. 2–5 and the references therein). Reinforcement Learning in Nephrology Although there have been recent applications of machine learning in nephrology,6,7 to the best of the authors' knowledge, the application of reinforcement learning to nephrology has been primarily limited to optimizing the erythropoietin dosage in hemodialysis patients.8,9 However, there are other settings where reinforcement learning has the potential to improve patient care in nephrology. For example, reinforcement learning methods can be adopted in the treatment of the complications of AKI or CKD (Figure 1). In this problem, the state models the conditions of the patient (e.g., vital signs, laboratory test results including urine and blood tests, and urine output measurements). The action refers to the treatment options (e.g., the dosage of medications such as sodium polystyrene sulfonate, and hemodialysis). The reward models the improvement in patient conditions. Similarly, reinforcement learning can help automate and optimize the dosage of immunosuppressive drugs in kidney transplants. Challenges and Opportunities Despite the success of reinforcement learning in several simplified clinical settings, their large-scale application to patient care faces several open challenges. The complexity of human biology complicates modeling clinical decision making as a reinforcement learning problem. The state space in such settings is often enormous, which could make a purely computational approach infeasible. Moreover, modeling all potential objectives a priori as a reward function may not be feasible. To overcome these challenges and realize the potential of reinforcement learning, clinical insight can play a pivotal role. More specifically, restricting the state space to only include highly relevant clinical variables could greatly reduce the computational complexity. Furthermore, using inverse reinforcement learning,2 relevant reward functions can be learned from retrospective studies assuming the optimality of clinical decisions. Another critical challenge is addressing moral and ethical concerns. It is imperative to ensure that reinforcement learning methods do not cause harm to the patient. To this end, there exists a need for a thorough validation of such methods before their use in patient care. Hence, there is a need to go beyond retrospective studies that have been used for the proof of concept of most existing reinforcement learning methods in health care applications.2,3 The lessons learned from the success of reinforcement learning in other application areas (e.g., self-driving cars) can help navigate the path to realizing its potential in health care. Accessible open-source simulation environments that enable researchers to compare various approaches are essential to the field of reinforcement learning. OpenAI Gym is currently the leading toolkit containing a wide range of simulated environments, e.g., surgical robotics.10 The development of high-quality and reliable simulation environments for nephrology and other health care applications can facilitate the development and validation of reinforcement learning methods beyond limited retrospective studies. The adoption of methods validated in such simulation environments in actual clinical settings will require clinicians' oversight. Similar to how self-driving cars require a human driver to ensure collision avoidance, clinicians' oversight is critical to ensure the safety of the patients, especially in the early stages of the adoption of reinforcement learning methods. The data from clinicians' decisions (e.g., overruling the automated treatment recommendation) can be used to improve the reliability of autonomous systems over time and reduce the burden of clinicians' oversight.

Similar Papers
  • Conference Article

QMDP: DASH Adaptation using Queueing Theory within a Markov Decision Process

  • Jan 09, 2021
  • Kevin Gatimu +1
  • Supplementary Content

Sample-Efficient I-Projections for Robot Learning

  • Apr 19, 2021
  • TUbilio (Technical University of Darmstadt)
  • Oleg Arenz
  • Conference Article
  • Citations15

Cola-HRL: Continuous-Lattice Hierarchical Reinforcement Learning for Autonomous Driving

  • Oct 23, 2022
  • Lingping Gao +7
  • Conference Article
  • Citations1

Comparison of Reinforcement Learning Methods for Production Control in Discrete Manufacturing Systems

  • Jun 12, 2023
  • Lingxiang Yun +3
  • Research Article
  • Citations1

EpiCare: A Reinforcement Learning Benchmark for Dynamic Treatment Regimes.

  • Jan 01, 2024
  • Advances in neural information processing systems
  • Mason Hargrave +2
  • Research Article
  • Citations5

5G Network Slicing: Methods to Support Blockchain and Reinforcement Learning

  • Mar 24, 2022
  • Computational Intelligence and Neuroscience
  • Juan Hu +1
  • Research Article
  • Citations31

The reinforcement learning method for occupant behavior in building control: A review

  • Sep 02, 2020
  • Energy and Built Environment
  • Mengjie Han +4
  • Research Article
  • Citations5

A Reinforcement Learning Method Using Reward Acquisition Efficiency for POMDP Environments

  • Jan 01, 2008
  • Transactions of the Japanese Society for Artificial Intelligence
  • Hirokazu Kawai +2
  • Conference Article
  • Citations5

Cohesion-driven Online Actor-Critic Reinforcement Learning for mHealth Intervention

  • Aug 15, 2018
  • Feiyun Zhu +4
  • Book Chapter
  • Citations11

Approximate Dynamic Programming

  • Feb 28, 2013
  • Rémi Munos
  • Research Article
  • Citations18

Autonomous UAV Maneuvering Decisions by Refining Opponent Strategies

  • Jun 01, 2024
  • IEEE Transactions on Aerospace and Electronic Systems
  • Like Sun +3
  • Research Article
  • Citations194

Deep reinforcement learning based energy management for a hybrid electric vehicle

  • Apr 14, 2020
  • Energy
  • Guodong Du +5
  • Conference Article
  • Citations23

Meta Reinforcement Learning-Based Lane Change Strategy for Autonomous Vehicles

  • Jul 11, 2021
  • Fei Ye +3
  • PDF
  • Research Article
  • Citations27

A new solution to distributed permutation flow shop scheduling problem based on NASH Q-Learning

  • Sep 30, 2021
  • Advances in Production Engineering & Management
  • J.F Ren +2
  • Research Article
  • Citations7

Learning to Acquire Whole-Body Humanoid Center of Mass Movements to Achieve Dynamic Tasks

  • Jan 01, 2008
  • Advanced Robotics
  • Takamitsu Matsubara +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.