• Home
  • Search
  • Importance sampling policy gradient algorithms in reproducing kernel Hilbert space
  • Cite Icon5
  • https://doi.org/10.1007/s10462-017-9579-xCopy DOI Icon

Importance sampling policy gradient algorithms in reproducing kernel Hilbert space

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Modeling policies in reproducing kernel Hilbert space (RKHS) offers a very flexible and powerful new family of policy gradient algorithms called RKHS policy gradient algorithms. They are designed to optimize over a space of very high or infinite dimensional policies. As a matter of fact, they are known to suffer from a large variance problem. This critical issue comes from the fact that updating the current policy is based on a functional gradient that does not exploit all old episodes sampled by previous policies. In this paper, we introduce a generalized RKHS policy gradient algorithm that integrates the following important ideas: (i) policy modeling in RKHS; (ii) normalized importance sampling, which helps reduce the estimation variance by reusing previously sampled episodes in a principled way; and (iii) regularization terms, which avoid updating the policy too over-fit to sampled data. In the experiment section, we provide an analysis of the proposed algorithms through bench-marking domains. The experiment results show that the proposed algorithm can still enjoy a powerful policy modeling in RKHS and achieve more data-efficiency.

Similar Papers
  • Conference Article

M-Power regularized least squares regression

  • May 01, 2017
  • Julien Audiffren +1
  • Research Article
  • Citations81

Gradient-Based Kernel Dimension Reduction for Regression

  • Jan 02, 2014
  • Journal of the American Statistical Association
  • Kenji Fukumizu +1
  • Conference Article
  • Citations106

Scalable Kernel Methods via Doubly Stochastic Gradients

  • Dec 08, 2014
  • neural information processing systems
  • Bo Dai +6
  • Research Article
  • Citations59

Laplacian least squares twin support vector machine for semi-supervised classification

  • May 28, 2014
  • Neurocomputing
  • Wei-Jie Chen +3
  • Research Article
  • Citations653

The Kernel Least-Mean-Square Algorithm

  • Feb 01, 2008
  • IEEE Transactions on Signal Processing
  • Weifeng Liu +2
  • Research Article
  • Citations1

Constrained approximate optimal transport maps

  • Jun 24, 2025
  • ESAIM: Control, Optimisation and Calculus of Variations
  • Eloi Tanguy +2
  • Research Article
  • Citations20

A Kernel-Based Indicator for Multi/Many-Objective Optimization

  • Aug 01, 2022
  • IEEE Transactions on Evolutionary Computation
  • Xinye Cai +6
  • Research Article
  • Citations11

Active learning based on minimization of the expected path-length of random walks on the learned manifold structure

  • Jun 06, 2017
  • Pattern Recognition
  • Chin-Chun Chang +1
  • Dissertation

Learning Invariances for High-Dimensional Data Analysis

  • Jan 01, 2014
  • The University of Queensland
  • Mahsa Baktashmotlagh
  • Research Article
  • Citations18

Reproducing kernel Hilbert spaces and variable metric algorithms in PDE-constrained shape optimization

  • May 03, 2017
  • Optimization Methods and Software
  • M Eigel +1
  • Research Article

Reproducing Kernel-Based Semiparametric Functional Smoothed Score Estimation with Binary Responses

  • Dec 11, 2025
  • Journal of Computational and Graphical Statistics
  • Meichen Liu +5
  • Research Article

Closed-Loop Control of Droplet Quality Based on Curriculum Deep Deterministic Policy Gradient Algorithm

  • Jan 01, 2025
  • IEEE Access
  • Yunyun Shen +5
  • Research Article
  • Citations27

The design and implementation of a deep reinforcement learning and quantum finance theory-inspired portfolio investment management system

  • Oct 25, 2023
  • Expert Systems With Applications
  • Yitao Qiu +2
  • PDF
  • Research Article
  • Citations3

A Policy Gradient Algorithm to Alleviate the Multi-Agent Value Overestimation Problem in Complex Environments

  • Nov 30, 2023
  • Sensors (Basel, Switzerland)
  • Yang Yang +4
  • Research Article

Deep Learning-Based Optimization for Mobile Robotic Delivery Systems

  • Nov 10, 2024
  • Optimizations in Applied Machine Learning
  • Diwei Zhu +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.