• Cite Icon5
  • https://doi.org/10.48550/arxiv.1901.04723Copy DOI Icon

Imitation-Regularized Offline Learning

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We study the problem of offline learning in automated decision systems under the contextual bandits model. We are given logged historical data consisting of contexts, (randomized) actions, and (nonnegative) rewards. A common goal is to evaluate what would happen if different actions were taken in the same contexts, so as to optimize the action policies accordingly. The typical approach to this problem, inverse probability weighted estimation (IPWE) [Bottou et al., 2013], requires logged action probabilities, which may be missing in practice due to engineering complications. Even when available, small action probabilities cause large uncertainty in IPWE, rendering the corresponding results insignificant. To solve both problems, we show how one can use policy improvement (PIL) objectives, regularized by policy imitation (IML). We motivate and analyze PIL as an extension to Clipped-IPWE, by showing that both are lower-bound surrogates to the vanilla IPWE. We also formally connect IML to IPWE variance estimation [Swaminathan and Joachims 2015] and natural policy gradients. Without probability logging, our PIL-IML interpretations justify and improve, by reward-weighting, the state-of-art cross-entropy (CE) loss that predicts the action items among all action candidates available in the same contexts. With probability logging, our main theoretical contribution connects IML-underfitting to the existence of either confounding variables or model misspecification. We show the value and accuracy of our insights by simulations based on Simpson's paradox, standard UCI multiclass-to-bandit conversions and on the Criteo counterfactual analysis challenge dataset.

Similar Papers
  • Research Article
  • Citations1

Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate

  • May 14, 2025
  • Operations Research
  • Yifan Lin +2
  • Book Chapter
  • Citations16

A New Natural Policy Gradient by Stationary Distribution Metric

  • Jan 01, 2008
  • Tetsuro Morimura +3
  • Book Chapter
  • Citations343

Natural Actor-Critic

  • Jan 01, 2005
  • Jan Peters +2
  • Conference Article
  • Citations4

Quasi-Newton Iteration in Deterministic Policy Gradient

  • Jun 08, 2022
  • Arash Bahari Kordabad +3
  • Research Article

Dilated Balanced cross entropy loss for medical image segmentation.

  • Feb 25, 2026
  • BMC medical imaging
  • Seyed Mohsen Hosseini +1
  • Dissertation
  • Citations1

Natural language processing as autoregressive generation

  • Jan 01, 2023
  • Xiang Lin
  • Conference Article
  • Citations3

Using visual speech information and perceptually motivated loss functions for binary mask estimation

  • Aug 25, 2017
  • Danny Websdale +1
  • Video Transcripts

A Study of Syntactic Multi-Modality in Non-Autoregressive Machine Translation

  • Jun 27, 2022
  • Underline Science Inc.
  • Kexun Zhang +6
  • Research Article
  • Citations61

Semi-Supervised Contrastive Learning With Similarity Co-Calibration

  • Jan 01, 2023
  • IEEE Transactions on Multimedia
  • Yuhang Zhang +5
  • Research Article
  • Citations31

Learning pseudo labels for semi-and-weakly supervised semantic segmentation

  • Jul 23, 2022
  • Pattern Recognition
  • Yude Wang +3
  • Book Chapter
  • Citations7

Benchmarking the Natural Gradient in Policy Gradient Methods and Evolution Strategies

  • Jan 01, 2021
  • Kay Hansel +2
  • Research Article
  • Citations20

Non-uniform Label Smoothing for Diabetic Retinopathy Grading from Retinal Fundus Images with Deep Neural Networks.

  • Jun 30, 2020
  • Translational Vision Science & Technology
  • Adrian Galdran +6
  • Research Article
  • Citations2

Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental Learning

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Zijian Gao +7
  • Conference Article
  • Citations22

Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training

  • Jun 06, 2021
  • Monisankha Pal +5
  • Conference Article
  • Citations11

Loss functions for optimizing Kappa as the evaluation measure for classifying diabetic retinopathy and prostate cancer images

  • Nov 26, 2020
  • Rajendran Nirthika +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.