• Home
  • Search
  • Strategic Exploration in Reinforcement Learning - New Algorithms and Learning Guarantees
  • Cite Icon6
  • https://doi.org/10.1184/r1/11882796.v1Copy DOI Icon

Strategic Exploration in Reinforcement Learning - New Algorithms and Learning Guarantees

Show More
  • Abstract
  • Literature Map
  • Citations
  • Similar Papers
Abstract

Reinforcement learning (RL) focuses on an essential aspect of intelligent behavior – how an agent can learn to make good decisions given experience and rewards in a stochastic<br>world. Yet popular RL algorithms that have enabled exciting successes in domains with good simulators (Go, Atari, etc) still often fail to learn in other domains because they rely on<br>simple heuristics for exploration. This provides additional empirical justification for essential questions around RL, specifically around algorithms that learn in a provably efficient manner through strategic exploration in any considered domain. This thesis provides new algorithms<br>and theory that enable good performance with respect to existing theoretical frameworks for evaluating RL algorithms (specifically, probably approximately correct) and introduces new stronger evaluation criteria, that may be particularly of interest as RL is applied to more real world problems.<br>For the first line of work on probably approximately correct (PAC) RL algorithms, we introduce a series of algorithms for episodic tabular domains with substantially better PAC<br>sample complexity bounds that culminate in a new algorithm with close to minimax optimal PAC and regret bounds. Look up tables are required by most sample efficient and computationally tractable algorithms, but cannot represent many practical domains. We therefore also present a new RL algorithm that can learn a good policy in environments with high dimensional observations and hidden deterministic states; unlike predecessors, this algorithm provably<br>explores not only in a statistically but also computationally efficient manner assuming access to function classes with efficient optimization oracles. To make progress it is critical to have the right measures of success. While empirical<br>demonstrations are quite clear, we find that for theoretical properties, two of the most commonly used learning frameworks, PAC guarantees and regret guarantees, each allow undesirable algorithm behavior (e.g. ignoring new observations that could improve the policy). We present<br>a new stronger learning framework called Uniform-PAC that unifies the existing frameworks and prevents undesirable algorithm properties. One caveat of all existing learning frameworks is that for any particular episode, we do not<br>know how well the algorithm will perform. To address this, we introduce the IPOC framework that requires algorithms to provide a certificate before each episode bounding how suboptimal the current policy can be. Such certifications may be of substantial interest in high stakes scenarios when an organization may wish to track or even pause an online RL system should the potential expected performance bound drop below a required expected outcome.

Similar Papers
  • Research Article
  • Citations26

Probably Approximately Correct (PAC) exploration in reinforcement learning

  • Jan 01, 2007
  • Rutgers University Community Repository (Rutgers University)
  • Alexander L Strehl
  • PDF
  • Research Article
  • Citations8

Long-Term Visitation Value for Deep Exploration in Sparse-Reward Reinforcement Learning

  • Feb 28, 2022
  • Algorithms
  • Simone Parisi +5
  • Conference Article

On the Design of Safe Continual RL Methods for Control of Nonlinear Systems

  • Jun 24, 2025
  • Austin Coursey +2
  • Research Article
  • Citations5

P2P power trading based on reinforcement learning for nanogrid clusters

  • Jul 19, 2024
  • Expert Systems With Applications
  • Hojun Jin +4
  • Supplementary Content

Bidirectional Human-Robot Learning: Imitation and Skill Improvement

  • Jun 23, 2020
  • TUbilio (Technical University of Darmstadt)
  • Sousa Ewerton +1
  • Research Article
  • Citations1

A Unified Analysis of Value-Function-Based Reinforcement Learning Algorithms

  • Nov 01, 1999
  • Neural Computation
  • Szepesváricsaba +1
  • PDF
  • Research Article
  • Citations14

Secure State Estimation of Cyber-Physical System under Cyber Attacks: Q-Learning vs. SARSA

  • Oct 01, 2022
  • Electronics
  • Zengwang Jin +5
  • Conference Article
  • Citations7

Constrained Expectation-Maximization Methods for Effective Reinforcement Learning

  • Jul 01, 2018
  • Gang Chen +2
  • Research Article

Stability and Convergence Analysis of Reinforcement Learning Algorithms in Complex Environments

  • Aug 21, 2025
  • Advances in Computer and Communication
  • Jifan Zhang
  • PDF
  • Research Article
  • Citations10

A Novel Functional Electrical Stimulation-Induced Cycling Controller Using Reinforcement Learning to Optimize Online Muscle Activation Pattern

  • Nov 24, 2022
  • Sensors
  • Tiago Coelho-Magalhães +2
  • Conference Article

Effective Linear Policy Gradient Search through Primal-Dual Approximation

  • Jul 01, 2020
  • Yiming Peng +2
  • Conference Article
  • Citations22

Reinforcement Learning For Field Development Policy Optimization

  • Oct 19, 2020
  • Giorgio De Paola +3
  • Research Article

Multi‐Agent Reinforcement Learning Algorithm Based on Local Observation Imitation Learning

  • Jan 01, 2025
  • IET Control Theory &amp; Applications
  • Hui Zhang +3
  • Conference Article
  • Citations5

Towards High-Level Intrinsic Exploration in Reinforcement Learning

  • Jul 01, 2020
  • Nicolas Bougie +1
  • Research Article
  • Citations5

Deep Deterministic Policy Gradient to Regulate Feedback Control Systems Using Reinforcement Learning

  • Jan 01, 2022
  • Computers, Materials &amp; Continua
  • Samir Salem Al-Bawri +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.