• Home
  • Search
  • PUMA: Performance Unchanged Model Augmentation for Training Data Removal
  • Cite Icon43
  • https://doi.org/10.1609/aaai.v36i8.20846Copy DOI Icon

PUMA: Performance Unchanged Model Augmentation for Training Data Removal

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Preserving the performance of a trained model while removing unique characteristics of marked training data points is challenging. Recent research usually suggests retraining a model from scratch with remaining training data or refining the model by reverting the model optimization on the marked data points. Unfortunately, aside from their computational inefficiency, those approaches inevitably hurt the resulting model's generalization ability since they remove not only unique characteristics but also discard shared (and possibly contributive) information. To address the performance degradation problem, this paper presents a novel approach called Performance Unchanged Model Augmentation (PUMA). The proposed PUMA framework explicitly models the influence of each training data point on the model's generalization ability with respect to various performance criteria. It then complements the negative impact of removing marked data by reweighting the remaining data optimally. To demonstrate the effectiveness of the PUMA framework, we compared it with multiple state-of-the-art data removal techniques in the experiments, where we show the PUMA can effectively and efficiently remove the unique characteristics of marked training data without retraining the model that can 1) fool a membership attack, and 2) resist performance degradation. In addition, as PUMA estimates the data importance during its operation, we show it could serve to debug mislabelled data points more efficiently than existing approaches.

Similar Papers
  • Research Article
  • Citations87

Robust hashing with local models for approximate similarity search.

  • Jul 01, 2014
  • IEEE Transactions on Cybernetics
  • Jingkuan Song +4
  • Conference Article

Learning prototype-based classifiers by margin maximization

  • Jun 01, 2017
  • Chiharu Wakou +2
  • Research Article
  • Citations573

Learning k for kNN Classification

  • Jan 12, 2017
  • ACM Transactions on Intelligent Systems and Technology
  • Shichao Zhang +4
  • Peer Review Report

Editor's evaluation: Robust and Efficient Assessment of Potency (REAP) as a quantitative tool for dose-response curve estimation

  • May 09, 2022
  • Philip Boonstra
  • Research Article
  • Citations59

Functions of Multiple Instances for Learning Target Signatures

  • Aug 01, 2015
  • IEEE Transactions on Geoscience and Remote Sensing
  • Changzhe Jiao +1
  • Research Article

Improving Text Classification by Leveraging Large Language Models for Data Augmentation

  • Jan 01, 2024
  • Academic Journal of Computing & Information Science
  • Siyun Yu
  • Research Article
  • Citations5

Detecting Source Contextual Barriers for Understanding Neural Machine Translation

  • Jan 01, 2021
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Guanlin Li +5
  • Book Chapter
  • Citations1

Building Weighted Classifier Ensembles Through Classifiers Pruning

  • Jan 01, 2018
  • Chenwei Cai +3
  • Conference Article
  • Citations74

Automated contour mapping using triangular element data structures and an interpolant over each irregular triangular domain

  • Jul 20, 1977
  • C M Gold +2
  • Conference Article
  • Citations1

Domain Generalization Via Adversarially Learned Novel Domains

  • Jul 18, 2022
  • Yu Zhe +3
  • Research Article
  • Citations24

Top-of-Mind Awareness and Share of Families: An Observation

  • May 01, 1969
  • Journal of Marketing Research
  • Alin Gruber
  • Research Article

Infrared Image Quality Estimation with Node-to-Graph Regression

  • Jan 01, 2026
  • IEEE Transactions on Multimedia
  • Ke Gu +8
  • Research Article

Frequency-aware domain randomization for single-source domain generalization in medical image segmentation.

  • Nov 29, 2025
  • Medical physics
  • Jiayi Wu +6
  • Book Chapter

CHEAPS2AGA: Bounding Space Usage in Variance-Reduced Stochastic Gradient Descent over Streaming Data and Its Asynchronous Parallel Variants

  • Jan 01, 2020
  • Yaqiong Peng +4
  • PDF
  • Research Article
  • Citations1

Applying Behavioral Biometrics to Mobile Device Use Measurement in Children: Evaluating the Impact of Training Data Size, Proximity, and Type on Model Performance.

  • Jul 07, 2025
  • Journal of technology in behavioral science
  • Olivia L Finnegan +13
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.