- Research Article
8
- 10.2139/ssrn.3980837
Ambiguous Dynamic Treatment Regimes: A Reinforcement Learning Approach
- Jan 01, 2021
- SSRN Electronic Journal
- Soroush Saghafian
Ambiguous Dynamic Treatment Regimes: A Reinforcement Learning Approach
While optimal dynamic treatment regimes (DTRs) can be estimated without specification of a predictive model, a model-based approach, combined with dynamic programming and Monte Carlo integration, enables direct probabilistic comparisons between the outcomes under the optimal DTR and alternative (dynamic or static) treatment regimes. The Bayesian predictive approach also circumvents problems related to frequentist estimators under the nonregular estimation problem. However, the model-based approach is susceptible to misspecification, in particular of the "null-paradox" type, which is due to the model parameters not having a direct causal interpretation in the presence of latent individual-level characteristics. Because it is reasonable to insist on correct inferences under the null of no difference between the alternative treatment regimes, we discuss how to achieve this through a "null-robust" reparametrization of the problem in a longitudinal setting. Since we argue that causal inference can be entirely understood as posterior predictive inference in a hypothetical population without covariate imbalances, we also discuss how controlling for confounding through inverse probability of treatment weighting can be justified and incorporated in the Bayesian setting.
Ambiguous Dynamic Treatment Regimes: A Reinforcement Learning Approach
Ambiguous Dynamic Treatment Regimes: A Reinforcement Learning Approach
Bayesian set of best dynamic treatment regimes: Construction and sample size calculation for SMARTswithbinary outcomes.
Sequential, multiple assignment, randomized trials (SMARTs) compare sequences of treatment decision rules called dynamic treatment regimes (DTRs). In particular, the Adaptive Treatment for Alcohol and Cocaine Dependence (ENGAGE) SMART aimed to determine the best DTRs for patients with a substance use disorder. While many authors have focused on a single pairwise comparison, addressing the main goal involves comparisons of >2 DTRs. For complex comparisons, there is a paucity of methods for binary outcomes. We fill this gap by extending the multiple comparisons with the best (MCB) methodology to the Bayesian binary outcome setting. The set of best is constructed based on simultaneous credible intervals. A substantial challenge for power analysis is the correlation between outcome estimators for distinct DTRs embedded in SMARTs due to overlapping subjects. We address this using Robins' G-computation formula to take a weighted average of parameter draws obtained via simulation from the parameter posteriors. We use non-informative priors and work with the exact distribution of parameters avoiding unnecessary normality assumptions and specification of the correlation matrix of DTR outcome summary statistics. We conduct simulation studies for both the construction of a set of optimal DTRs using the Bayesian MCB procedure and the sample size calculation for two common SMART designs. We illustrate our method on the ENGAGE SMART. The R package SMARTbayesR for power calculations is freely available on the Comprehensive R Archive Network (CRAN) repository. An RShiny app is available at https://wilart.shinyapps.io/shinysmartbayesr/.
Read moreChronic Kidney Disease-Mineral and Bone Disorder Management in 4D: The Case for Dynamic Treatment Regime Methods to Optimize Care
Purpose of ReviewChronic Kidney Disease-Mineral and Bone Disorder (CKD-MBD) is a complex condition impacting patients with kidney failure and characterized by inter-related features such as hyperparathyroidism, hyperphosphatemia, and hypocalcemia. Current treatments include active vitamin D sterols, calcimimetics, and phosphate binders alone and in combination. However, identifying optimal treatment is challenged by interdependency among CKD-MBD features, requiring new approaches to understand dynamic systems. In this review, we discuss challenges and opportunities for a more integrated view of CKD-MBD care.Recent FindingsFew clinical studies in CKD-MBD care have incorporated a dynamic understanding of the disorder and its treatment. Dynamic treatment regime methods are an evolving area of artificial intelligence (AI) that offer a promising approach for modeling and understanding CKD-MBD care. Efforts to date have included dynamic systems and quantitative systems pharmacology-based models to simulate the impact of alternative treatment regimes. Additional studies utilizing dynamic treatment regime approaches may help improve knowledge gaps in CKD-MBD care.SummaryAlthough preliminary research highlights the potential of dynamic treatment regime approaches in optimizing CKD-MBD management, further investigation and clinical validation are necessary to fully harness this approach for improving patient outcomes.
Read moreReinforcement learning in clinical medicine: a method to optimize dynamic treatment regime over time.
Precision medicine requires individualized treatment regime for subjects with different clinical characteristics. Machine learning methods have witnessed rapid progress in recent years, which can be employed to make individualized treatment regime in clinical practice. The idea of reinforcement learning method is to take action in response to the changing environment. In clinical medicine, this idea can be used to assign optimal regime to patients with distinct characteristics. In the field of statistics, reinforcement learning has been widely investigated, aiming to identify an optimal dynamic treatment regime (DTR). Q-learning is among the earliest methods to identify optimal DTR, which fits linear outcome models in a recursive manner. The advantage is its easy interpretation and can be performed in most statistical software. However, it suffers from the risk of misspecification of the linear model. More recently, some other methods not so heavily depend on model specification have been developed such as inverse probability weighted estimator and augmented inverse probability weighted estimator. This review introduces the basic ideas of these methods and shows how to perform the learning algorithm within R environment.
Read moreAdversarial Cooperative Imitation Learning for Dynamic Treatment Regimes✱
Recent developments in discovering dynamic treatment regimes (DTRs) have heightened the importance of deep reinforcement learning (DRL) which are used to recover the doctor’s treatment policies. However, existing DRL-based methods expose the following limitations: 1) supervised methods based on behavior cloning suffer from compounding errors; 2) the self-defined reward signals in reinforcement learning models are either too sparse or need clinical guidance; 3) only positive trajectories (e.g. survived patients) are considered in current imitation learning models, with negative trajectories (e.g. deceased patients) been largely ignored, which are examples of what not to do and could help the learned policy avoid repeating mistakes. To address these limitations, in this paper, we propose the adversarial cooperative imitation learning model, ACIL, to deduce the optimal dynamic treatment regimes that mimics the positive trajectories while differs from the negative trajectories. Specifically, two discriminators are used to help achieve this goal: an adversarial discriminator is designed to minimize the discrepancies between the trajectories generated from the policy and the positive trajectories, and a cooperative discriminator is used to distinguish the negative trajectories from the positive and generated trajectories. The reward signals from the discriminators are utilized to refine the policy for dynamic treatment regimes. Experiments on the publicly real-world medical data demonstrate that ACIL improves the likelihood of patient survival and provides better dynamic treatment regimes with the exploitation of information from both positive and negative trajectories.
Read moreInference for optimal dynamic treatment regimes using an adaptive m-out-of-n bootstrap scheme.
A dynamic treatment regime consists of a set of decision rules that dictate how to individualize treatment to patients based on available treatment and covariate history. A common method for estimating an optimal dynamic treatment regime from data is Q-learning which involves nonsmooth operations of the data. This nonsmoothness causes standard asymptotic approaches for inference like the bootstrap or Taylor series arguments to breakdown if applied without correction. Here, we consider the m-out-of-n bootstrap for constructing confidence intervals for the parameters indexing the optimal dynamic regime. We propose an adaptive choice of m and show that it produces asymptotically correct confidence sets under fixed alternatives. Furthermore, the proposed method has the advantage of being conceptually and computationally much simple than competing methods possessing this same theoretical property. We provide an extensive simulation study to compare the proposed method with currently available inference procedures. The results suggest that the proposed method delivers nominal coverage while being less conservative than alternatives. The proposed methods are implemented in the qLearn R-package and have been made available on the Comprehensive R-Archive Network (http://cran.r-project.org/). Analysis of the Sequenced Treatment Alternatives to Relieve Depression (STAR*D) study is used as an illustrative example.
Read moreComment on "Dynamic treatment regimes: technical challenges and applications"
Inference for parameters associated with optimal dynamic treatment regimes is challenging as these estimators are nonregular when there are non-responders to treatments. In this discussion, we comment on three aspects of alleviating this nonregularity. We first discuss an alternative approach for smoothing the quality functions. We then discuss some further details on our existing work to identify non-responders through penalization. Third, we propose a clinically meaningful value assessment whose estimator does not suffer from nonregularity.
Read moreComparison of Dynamic Treatment Regimes via Inverse Probability Weighting
Appropriate analysis of observational data is our best chance to obtain answers to many questions that involve dynamic treatment regimes. This paper describes a simple method to compare dynamic treatment regimes by artificially censoring subjects and then using inverse probability weighting (IPW) to adjust for any selection bias introduced by the artificial censoring. The basic strategy can be summarized in four steps: 1) define two regimes of interest, 2) artificially censor individuals when they stop following one of the regimes of interest, 3) estimate inverse probability weights to adjust for the potential selection bias introduced by censoring in the previous step, 4) compare the survival of the uncensored individuals under each regime of interest by fitting an inverse probability weighted Cox proportional hazards model with the dichotomous regime indicator and the baseline confounders as covariates. In the absence of model misspecification, the method is valid provided data are available on all time-varying and baseline joint predictors of survival and regime discontinuation. We present an application of the method to compare the AIDS-free survival under two dynamic treatment regimes in a large prospective study of HIV-infected patients. The paper concludes by discussing the relative advantages and disadvantages of censoring/IPW versus g-estimation of nested structural models to compare dynamic regimes.
Read moreWeighted Q-learning for optimal dynamic treatment regimes with nonignorable missing covariates
Dynamic treatment regimes (DTRs) formalize medical decision-making as a sequence of rules for different stages, mapping patient-level information to recommended treatments. In practice, estimating an optimal DTR using observational data from electronic medical record (EMR) databases can be complicated by nonignorable missing covariates resulting from informative monitoring of patients. Since complete case analysis can provide consistent estimation of outcome model parameters under the assumption of outcome-independent missingness, Q-learning is a natural approach to accommodating nonignorable missing covariates. However, the backward induction algorithm used in Q-learning can introduce challenges, as nonignorable missing covariates at later stages can result in nonignorable missing pseudo-outcomes at earlier stages, leading to suboptimal DTRs, even if the longitudinal outcome variables are fully observed. To address this unique missing data problem in DTR settings, we propose 2 weighted Q-learning approaches where inverse probability weights for missingness of the pseudo-outcomes are obtained through estimating equations with valid nonresponse instrumental variables or sensitivity analysis. The asymptotic properties of the weighted Q-learning estimators are derived, and the finite-sample performance of the proposed methods is evaluated and compared with alternative methods through extensive simulation studies. Using EMR data from the Medical Information Mart for Intensive Care database, we apply the proposed methods to investigate the optimal fluid strategy for sepsis patients in intensive care units.
Read moreEpiCare: A Reinforcement Learning Benchmark for Dynamic Treatment Regimes.
Healthcare applications pose significant challenges to existing reinforcement learning (RL) methods due to implementation risks, limited data availability, short treatment episodes, sparse rewards, partial observations, and heterogeneous treatment effects. Despite significant interest in using RL to generate dynamic treatment regimes for longitudinal patient care scenarios, no standardized benchmark has yet been developed. To fill this need we introduce Episodes of Care (EpiCare), a benchmark designed to mimic the challenges associated with applying RL to longitudinal healthcare settings. We leverage this benchmark to test five state-of-the-art offline RL models as well as five common off-policy evaluation (OPE) techniques. Our results suggest that while offline RL may be capable of improving upon existing standards of care given sufficient data, its applicability does not appear to extend to the moderate to low data regimes typical of current healthcare settings. Additionally, we demonstrate that several OPE techniques standard in the the medical RL literature fail to perform adequately on our benchmark. These results suggest that the performance of RL models in dynamic treatment regimes may be difficult to meaningfully evaluate using current OPE methods, indicating that RL for this application domain may still be in its early stages. We hope that these results along with the benchmark will facilitate better comparison of existing methods and inspire further research into techniques that increase the practical applicability of medical RL.
Read moreMultistage Treatment Regimes
For each patient, treatment for a disease often is a multistage process involving an alternating sequence of observations and therapeutic decisions, with the physician’s decision at each stage based on the patient’s entire history up to that stage. This chapter begins with discussion of a simple two-stage version of this process, in which a Frontline treatment is given initially and, if and when the patient’s disease worsens, i.e., progresses, a second, Salvage treatment is given, so the two-stage regime is (Frontline, Salvage). The discussion of this case will include examples where, if one only accounts for the effects of Frontline and Salvage separately in each stage, the effect of the entire regime on survival time may not be obvious. Discussions and illustrations will be given of the general paradigms of dynamic treatment regimes (DTRs) and sequential multiple assignment randomized trials (SMARTs). Several statistical analyses of data from a prostate cancer trial designed by Thall et al. (2000) to evaluate multiple DTRs then will be discussed in detail. As a final example, several statistical analyses of observational data from a semi-SMART design of DTRs for acute leukemia, given by Estey et al. (1999), Wahed and Thall (2013), and Xu et al. (2016), will be discussed.
Read moreScreening Experiments for Developing Dynamic Treatment Regimes
Dynamic treatment regimes are time-varying treatments that individualize sequences of treatments to the patient. The construction of dynamic treatment regimes is challenging because a patient will be eligible for some treatment components only if he has not responded (or has responded) to other treatment components. In addition, there are usually a number of potentially useful treatment components and combinations thereof. In this article, we propose new methodology for identifying promising components and screening out negligible ones. First, we define causal factorial effects for treatment components that may be applied sequentially to a patient. Second, we propose experimental designs that can be used to study the treatment components. Surprisingly, modifications can be made to (fractional) factorial designs—more commonly found in the engineering statistics literature—for screening in this setting. Furthermore, we provide an analysis model that can be used to screen the factorial effects. We demonstrate the proposed methodology using examples motivated in the literature and also via a simulation study.
Read moreChange-point detection for infinite horizon dynamic treatment regimes.
A dynamic treatment regime is a set of decision rules for how to treat a patient at multiple time points. At each time point, a treatment decision is made depending on the patient's medical history up to that point. We consider the infinite-horizon setting in which the number of decision points is very large. Specifically, we consider long trajectories of patients' measurements recorded over time. At each time point, the decision whether to intervene or not is conditional on whether or not there was a change in the patient's trajectory. We present change-point detection tools and show how to use them in defining dynamic treatment regimes. The performance of these regimes is assessed using an extensive simulation study. We demonstrate the utility of the proposed change-point detection approach using two case studies: detection of sepsis in preterm infants in the intensive care unit and detection of a change in glucose levels of a diabetic patient.
Read moreDynamic treatment regimes with interference
Precision medicine describes health care where patient‐level data are used to inform treatment decisions. Within this framework, dynamic treatment regimes (DTRs) are sequences of decision rules that take individual patient information as input data and then output treatment recommendations. DTR estimation from observational data typically relies on the assumption of no interference: i.e., the outcome of one individual is unaffected by the treatment assignment of others. However, in many social network contexts, such as friendship or family networks, and for many health concerns, such as infectious diseases, this assumption is questionable. We investigate the DTR estimation method of dynamic weighted ordinary least squares (dWOLS), which boasts of easy implementation and the so‐called double‐robustness property, but relies on the assumption of no interference. We define a network propensity function and build on it to establish an implementation of dWOLS that remains doubly robust under interference associated with network links. The method's properties are demonstrated via simulation and applied to data from the Population Assessment of Tobacco and Health (PATH) study to investigate cigarette dependence within two‐person household networks.
Read moreDynamic Treatment Regimes
Dynamic treatment regimes describe a class of treatments or interventions, often given sequentially, that are tailored to the characteristics of an individual at the time the treatment decision is made. In this article, we introduce dynamic treatment regimes and sequential multiple assignment randomized trial designs, which give rise to the data that can be used to estimate the impact of such personalized treatment strategies.
Read more