• Open Access IconOpen Access
  • Cite Icon343
  • https://doi.org/10.1007/11564096_29Copy DOI Icon

Natural Actor-Critic

  • Jan 1, 2005
  • Jan Peters +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This paper investigates a novel model-free reinforcement learning architecture, the Natural Actor-Critic. The actor updates are based on stochastic policy gradients employing Amari’s natural gradient approach, while the critic obtains both the natural policy gradient and additional parameters of a value function simultaneously by linear regression. We show that actor improvements with natural policy gradients are particularly appealing as these are independent of coordinate frame of the chosen policy representation, and can be estimated more efficiently than regular policy gradients. The critic makes use of a special basis function parameterization motivated by the policy-gradient compatible function approximation. We show that several well-known reinforcement learning methods such as the original Actor-Critic and Bradtke’s Linear Quadratic Q-Learning are in fact Natural Actor-Critic algorithms. Empirical evaluations illustrate the effectiveness of our techniques in comparison to previous methods, and also demonstrate their applicability for learning control on an anthropomorphic robot arm.

Similar Papers
  • Book Chapter
  • Citations16

A New Natural Policy Gradient by Stationary Distribution Metric

  • Jan 01, 2008
  • Tetsuro Morimura +3
  • Research Article
  • Citations1

Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate

  • May 14, 2025
  • Operations Research
  • Yifan Lin +2
  • Conference Article
  • Citations4

Quasi-Newton Iteration in Deterministic Policy Gradient

  • Jun 08, 2022
  • Arash Bahari Kordabad +3
  • Research Article
  • Citations3

Convergence and Optimality of Policy Gradient Methods in Weakly Smooth Settings

  • Jun 28, 2022
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Matthew S Zhang +2
  • Research Article
  • Citations6

Deterministic policy gradient algorithms for semi‐Markov decision processes

  • Oct 13, 2021
  • International Journal of Intelligent Systems
  • Ashkan Haji Hosseinloo +1
  • Conference Article
  • Citations9

EMBB and URLLC Service Multiplexing Based on Deep Reinforcement Learning in 5G and Beyond

  • Apr 10, 2022
  • Yi-Huai Hsu +1
  • Research Article
  • Citations26

Direct and indirect reinforcement learning

  • May 31, 2021
  • International Journal of Intelligent Systems
  • Yang Guan +6
  • Book Chapter
  • Citations7

Benchmarking the Natural Gradient in Policy Gradient Methods and Evolution Strategies

  • Jan 01, 2021
  • Kay Hansel +2
  • Research Article
  • Citations2

Military Decision Support with Actor and Critic Reinforcement Learning Agents

  • Feb 26, 2024
  • Defence Science Journal
  • Jungmok Ma
  • Research Article
  • Citations22

A deep recurrent Q network towards self‐adapting distributed microservice architecture

  • Nov 28, 2019
  • Software: Practice and Experience
  • Basel Magableh +1
  • Conference Article

Ant Colony Optimization with Policy Gradients and Replay

  • Jul 13, 2025
  • William Jardee +1
  • Research Article
  • Citations9

Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning.

  • May 01, 2025
  • IEEE transactions on pattern analysis and machine intelligence
  • Shangding Gu +6
  • Research Article

CNN-based Reinforcement Learning with Policy Gradient for Khmer Chess

  • Apr 30, 2025
  • Techno-Science Research Journal
  • Both Chan
  • PDF
  • Research Article

A Fisher–Rao Gradient Flow for Entropy-Regularised Markov Decision Processes in Polish Spaces

  • Aug 11, 2025
  • Foundations of Computational Mathematics
  • Bekzhan Kerimkulov +4
  • Conference Article
  • Citations26

Risk-sensitive reinforcement learning

  • Oct 15, 2020
  • Nelson Vadori +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.