• Open Access IconOpen Access
  • Cite Icon7
  • https://doi.org/10.1007/978-3-031-19849-6_20Copy DOI Icon

Automata Learning Meets Shielding

  • Jan 1, 2022
  • Martin Tappler +5 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Abstract Safety is still one of the major research challenges in reinforcement learning (RL). In this paper, we address the problem of how to avoid safety violations of RL agents during exploration in probabilistic and partially unknown environments. Our approach combines automata learning for Markov Decision Processes (MDPs) and shield synthesis in an iterative approach. Initially, the MDP representing the environment is unknown. The agent starts exploring the environment and collects traces. From the collected traces, we passively learn MDPs that abstractly represent the safety-relevant aspects of the environment. Given a learned MDP and a safety specification, we construct a shield. For each state-action pair within a learned MDP, the shield computes exact probabilities on how likely it is that executing the action results in violating the specification from the current state within the next k steps. After the shield is constructed, the shield is used during runtime and blocks any actions that induce a too large risk from the agent. The shielded agent continues to explore the environment and collects new data on the environment. Iteratively, we use the collected data to learn new MDPs with higher accuracy, resulting in turn in shields able to prevent more safety violations. We implemented our approach and present a detailed case study of a Q-learning agent exploring slippery Gridworlds. In our experiments, we show that as the agent explores more and more of the environment during training, the improved learned models lead to shields that are able to prevent many safety violations.KeywordsAutomata learningShieldingMarkov Decision Processes

Similar Papers
  • Research Article
  • Citations26

Probably Approximately Correct (PAC) exploration in reinforcement learning

  • Jan 01, 2007
  • Rutgers University Community Repository (Rutgers University)
  • Alexander L Strehl
  • Conference Article

Out-of-Distribution Detection for Reinforcement Learning Agents with Probabilistic Dynamics Models

  • May 30, 2023
  • Tom Haider +3
  • Research Article
  • Citations13

Reinforcement Learning for Clinical Applications.

  • Feb 08, 2023
  • Clinical Journal of the American Society of Nephrology
  • Kia Khezeli +5
  • Components
  • Citations7

Reward-predictive representations generalize across tasks in reinforcement learning

  • Oct 15, 2020
  • Lucas Lehnert +3
  • Book Chapter
  • Citations17

Solving Markov Decision Processes via Simulation

  • Sep 18, 2014
  • Abhijit Gosavi
  • Research Article

Don’t Follow Reinforcement Learning Blindly: Lower Sample Complexity of Learning Optimal Inventory Control Policies with Fixed Ordering Cost

  • Sep 03, 2025
  • Production and Operations Management
  • Xiaoyu Fan +4
  • Research Article
  • Citations28

Exploring reinforcement learning in process control: a comprehensive survey

  • Feb 28, 2025
  • International Journal of Systems Science
  • N Rajasekhar +2
  • Conference Instance
  • Citations21

Learning Curriculum Policies for Reinforcement Learning

  • May 08, 2019
  • Sanmit Narvekar +1
  • Conference Article
  • Citations1

Agent sensing with stateful resources

  • Nov 30, 2010
  • Adam Eck +1
  • Research Article
  • Citations10

The Unreasonable Effectiveness of Inverse Reinforcement Learning in Advancing Cancer Research

  • Apr 03, 2020
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • John Kalantari +2
  • Conference Article

Function approximation for large markov decision processes using self-organizing neural networks

  • Jul 01, 2015
  • Teck-Hou Teng
  • Research Article
  • Citations64

When blockchain meets AI: Optimal mining strategy achieved by machine learning

  • Feb 01, 2021
  • International Journal of Intelligent Systems
  • Taotao Wang +2
  • Book Chapter
  • Citations12

A Survey of Optimistic Planning in Markov Decision Processes

  • Dec 17, 2012
  • Lucian Buşoniu +2
  • Research Article
  • Citations306

Advanced planning for autonomous vehicles using reinforcement learning and deep inverse reinforcement learning

  • Jan 15, 2019
  • Robotics and Autonomous Systems
  • Changxi You +3
  • Conference Article
  • Citations3

Trading financial assets with actor critic using Kronecker-factored trust region (ACKTR)

  • Jan 01, 2020
  • AIP conference proceedings
  • F Heryanto +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.