• Home
  • Search
  • Corruption-Robust Exploration in Episodic Reinforcement Learning
  • Open Access IconOpen Access
  • Cite Icon48
  • https://doi.org/10.1287/moor.2021.0202Copy DOI Icon

Corruption-Robust Exploration in Episodic Reinforcement Learning

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We initiate the study of episodic reinforcement learning (RL) under adversarial corruptions in both the rewards and the transition probabilities of the underlying system, extending recent results for the special case of multiarmed bandits. We provide a framework that modifies the aggressive exploration enjoyed by existing reinforcement learning approaches based on optimism in the face of uncertainty by complementing them with principles from action elimination. Importantly, our framework circumvents the major challenges posed by naively applying action elimination in the RL setting, as formalized by a lower bound we demonstrate. Our framework yields efficient algorithms that (a) attain near-optimal regret in the absence of corruptions and (b) adapt to unknown levels of corruption, enjoying regret guarantees that degrade gracefully in the total corruption encountered. To showcase the generality of our approach, we derive results for both tabular settings (where states and actions are finite) and linear Markov decision process settings (where the dynamics and rewards admit a linear underlying representation). Notably, our work provides the first sublinear regret guarantee that accommodates any deviation from purely independent and identically distributed transitions in the bandit-feedback model for episodic reinforcement learning. Supplemental Material: The online appendix is available at https://doi.org/10.1287/moor.2021.0202 .

Similar Papers
  • Research Article

Logarithmic Regret for Linear Markov Decision Processes with Adversarial Corruptions

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Canzhe Zhao +3
  • Research Article

A Comparative Analysis of Self-Aware Reinforcement Learning Models for Real-Time Intrusion Detection in Fog Networks

  • Feb 14, 2026
  • Future Internet
  • Nyashadzashe Tamuka +5
  • Research Article
  • Citations99

The relation between reinforcement learning parameters and the influence of reinforcement history on choice behavior

  • May 15, 2015
  • Journal of Mathematical Psychology
  • Kentaro Katahira
  • PDF
  • Research Article
  • Citations18

A comparison of reinforcement learning models of human spatial navigation

  • Aug 17, 2022
  • Scientific Reports
  • Qiliang He +4
  • Research Article
  • Citations93

The statistical structures of reinforcement learning with asymmetric value updates

  • Oct 01, 2018
  • Journal of Mathematical Psychology
  • Kentaro Katahira
  • Conference Article
  • Citations5

Reinforcement Learning based Carbon Nanotube Growth Automation

  • Oct 12, 2021
  • Ashish Pandey +3
  • Book Chapter
  • Citations20

Active Finite Reward Automaton Inference and Reinforcement Learning Using Queries and Counterexamples

  • Jan 01, 2021
  • Zhe Xu +4
  • Research Article
  • Citations11

Data Quality, Bias, and Strategic Challenges in Reinforcement Learning for Healthcare: A Survey

  • Sep 20, 2024
  • International Journal of Data Informatics and Intelligent Computing
  • Atta Ur Rahman +3
  • Research Article
  • Citations6

A reinforcement learning and predictive analytics approach for enhancing credit assessment in manufacturing

  • Jun 01, 2025
  • Decision Analytics Journal
  • Abdul Razaque +4
  • Book Chapter

Effects of Stress and Genotype on Meta-parameter Dynamics in Reinforcement Learning

  • Sep 07, 2007
  • Gediminas Lukšys +4
  • Dissertation
  • Citations1

Acoustic Cloak Design Using Generative Modeling and Reinforcement Learning

  • Jan 01, 2022
  • Linwei Zhuo
  • Conference Article
  • Citations71

Learning from Demonstration for Shaping through Inverse Reinforcement Learning

  • May 09, 2016
  • Halit Bener Suay +3
  • Research Article
  • Citations2

Reinforcement Learning with Foregone Payoff Information in Normal Form Games

  • Mar 20, 2019
  • SSRN Electronic Journal
  • Naoki Funai
  • Book Chapter

Chapter 11 - Does the Dopaminergic Error Signal Act Like a Cached-Value Prediction Error?

  • Jan 01, 2018
  • Goal-Directed Decision Making
  • Melissa J Sharpe +1
  • Research Article
  • Citations1

A Deep Reinforcement Learning–Based Urban Traffic Control Model for Vehicle‐to‐Everything Ecosystem

  • Jan 01, 2025
  • Journal of Advanced Transportation
  • Lingyu Zheng +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.