• Home
  • Search
  • Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games
  • https://doi.org/10.65109/tdox4540Copy DOI Icon

Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games

  • May 3, 2021
  • Kenshi Abe +1 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Off-policy evaluation (OPE) is the problem of evaluating new policies using historical data obtained from a different policy. In the recent OPE context, most studies have focused on single-player cases, and not on multi-player cases. In this study, we propose OPE estimators constructed by the doubly robust and double reinforcement learning estimators in two-player zero-sum Markov games. The proposed estimators project exploitability that is often used as a metric for determining how close a policy profile (i.e., a tuple of policies) is to a Nash equilibrium in two-player zero-sum games. We prove the exploitability estimation error bounds for the proposed estimators. We then propose the methods to find the best candidate policy profile by selecting the policy profile that minimizes the estimated exploitability from a given policy profile class. We prove the regret bounds of the policy profiles selected by our methods. Finally, we demonstrate the effectiveness and performance of the proposed estimators through experiments.

Similar Papers
  • Research Article
  • Citations7

Solution concepts for games with ambiguous payoffs

  • May 28, 2015
  • Theory and Decision
  • Dorian Beauchêne
  • Book Chapter
  • Citations4

Learning Algorithms for Differential Games of Continuous-Time Systems

  • Jan 01, 2017
  • Derong Liu +4
  • Research Article
  • Citations9

Pure Strategy Equilibria in Symmetric Two-Player Zero-Sum Games

  • Feb 22, 2010
  • SSRN Electronic Journal
  • Peter Dürsch +2
  • Research Article
  • Citations12

Computing approximate pure Nash equilibria in congestion games

  • Jun 01, 2012
  • ACM SIGecom Exchanges
  • Ioannis Caragiannis +3
  • Research Article
  • Citations15

Uniqueness of the index for Nash equilibria of two-player games

  • Sep 17, 1997
  • Economic Theory
  • Srihari Govindan +1
  • Conference Article
  • Citations21

Inverse two-player zero-sum dynamic games

  • Nov 01, 2016
  • Dorian Tsai +2
  • Research Article
  • Citations114

Learning in Games by Random Sampling

  • May 01, 2001
  • Journal of Economic Theory
  • James W Friedman +1
  • Conference Article
  • Citations30

Double-oracle algorithm for computing an exact nash equilibrium in zero-sum extensive-form games

  • Jun 27, 2013
  • Branislav Bošanský +4
  • Research Article
  • Citations1

An undecidable statement regarding zero-sum games

  • Feb 22, 2024
  • Games and Economic Behavior
  • Mark Fey
  • Conference Article
  • Citations1

Fictitious Cross-Play: Learning Global Nash Equilibrium in Mixed Cooperative-Competitive Games

  • May 30, 2023
  • Zelai Xu +4
  • Supplementary Content

Efficient algorithms for computing approximate equilibria in bimatrix, polymatrix and Lipschitz games

  • Nov 28, 2016
  • University of Liverpool
  • Argyrios Deligkas
  • Conference Article

Solving Two-player Games with QBF Solvers in General Game Playing

  • May 06, 2024
  • Yifan He +2
  • PDF
  • Research Article
  • Citations14

Secure State Estimation of Cyber-Physical System under Cyber Attacks: Q-Learning vs. SARSA

  • Oct 01, 2022
  • Electronics
  • Zengwang Jin +5
  • Research Article

Model-free policy iteration optimal control of fuzzy systems via a two-player zero-sum game

  • Mar 21, 2025
  • International Journal of Systems Science
  • Yifan Deng +2
  • PDF
  • Conference Article
  • Citations6

On Reinforcement Learning for Turn-based Zero-sum Markov Games

  • Oct 18, 2020
  • Devavrat Shah +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.