• Home
  • Search
  • Learning from Demonstration for Shaping through Inverse Reinforcement Learning
  • Cite Icon71
  • https://doi.org/10.5555/2936924.2936988Copy DOI Icon

Learning from Demonstration for Shaping through Inverse Reinforcement Learning

  • May 9, 2016
  • Halit Bener Suay +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Model-free episodic reinforcement learning problems define the environment reward with functions that often provide only sparse information throughout the task. Consequently, agents are not given enough feedback about the fitness of their actions until the task ends with success or failure. Previous work addresses this problem with reward shaping. In this paper we introduce a novel approach to improve model-free reinforcement learning agents' performance with a three step approach. Specifically, we collect demonstration data, use the data to recover a linear function using inverse reinforcement learning and we use the recovered function for potential-based reward shaping. Our approach is model-free and scalable to high dimensional domains. To show the scalability of our approach we present two sets of experiments in a two dimensional Maze domain, and the 27 dimensional Mario AI domain. We compare the performance of our algorithm to previously introduced reinforcement learning from demonstration algorithms. Our experiments show that our approach outperforms the state-of-the-art in cumulative reward, learning rate and asymptotic performance.

Similar Papers
  • Supplementary Content

Sample-Efficient I-Projections for Robot Learning

  • Apr 19, 2021
  • TUbilio (Technical University of Darmstadt)
  • Oleg Arenz
  • Peer Review Report

Reviewer #2 (Public review): Neural signatures of model-based and model-free reinforcement learning across prefrontal cortex and striatum

  • Feb 27, 2026
  • Bruno Miranda +5
  • Conference Article

Using reward/utility based impact scores in partitioning

  • May 05, 2014
  • William J Curran +2
  • Conference Article
  • Citations2

Context-dependent meta-control for reinforcement learning using a Dirichlet process Gaussian mixture model

  • Jan 01, 2018
  • Dongjae Kim +1
  • Research Article
  • Citations1

Perceptions of Foreigners about Process of Learning Turkish (Türkçe Öğrenen Yabancıların Öğrenme Süreçlerine Yönelik Algıları).....Doi: 10.14686/BUEFAD.201428189

  • Oct 31, 2014
  • Deniz Melanlıoğlu
  • PDF
  • Research Article
  • Citations2

Model-free inverse reinforcement learning with multi-intention, unlabeled, and overlapping demonstrations

  • Nov 30, 2022
  • Machine Learning
  • Ariyan Bighashdel +2
  • Conference Article
  • Citations15

Gradient-based inverse risk-sensitive reinforcement learning

  • Dec 01, 2017
  • Eric Mazumdar +3
  • PDF
  • Research Article
  • Citations108

A survey of inverse reinforcement learning

  • Feb 08, 2022
  • Artificial Intelligence Review
  • Stephen Adams +2
  • Research Article
  • Citations13

Reinforcement Learning for Clinical Applications.

  • Feb 08, 2023
  • Clinical Journal of the American Society of Nephrology
  • Kia Khezeli +5
  • Supplementary Content

Stability Control of Biped Robot on Static and Dynamic Platforms Based on Hybrid Reinforcement Learning

  • Dec 09, 2020
  • Figshare
  • Ao Xi
  • Book Chapter

Estimation of the Change of Agents Behavior Strategy Using State-Action History

  • Jan 01, 2017
  • Shihori Uchida +2
  • Research Article
  • Citations39

Navigating complex decision spaces: Problems and paradigms in sequential choice.

  • Jan 01, 2014
  • Psychological Bulletin
  • Matthew M Walsh +1
  • Research Article
  • Citations306

Advanced planning for autonomous vehicles using reinforcement learning and deep inverse reinforcement learning

  • Jan 15, 2019
  • Robotics and Autonomous Systems
  • Changxi You +3
  • PDF
  • Research Article
  • Citations6

Advances and applications in inverse reinforcement learning: a comprehensive review

  • Mar 26, 2025
  • Neural Computing and Applications
  • Saurabh Deshpande +4
  • Research Article
  • Citations26

Probably Approximately Correct (PAC) exploration in reinforcement learning

  • Jan 01, 2007
  • Rutgers University Community Repository (Rutgers University)
  • Alexander L Strehl
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.