• Home
  • Search
  • Quantifying and Analyzing Entity-Level Memorization in Large Language Models
  • Cite Icon4
  • https://doi.org/10.1609/aaai.v38i17.29948Copy DOI Icon

Quantifying and Analyzing Entity-Level Memorization in Large Language Models

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Large language models (LLMs) have been proven capable of memorizing their training data, which can be extracted through specifically designed prompts. As the scale of datasets continues to grow, privacy risks arising from memorization have attracted increasing attention. Quantifying language model memorization helps evaluate potential privacy risks. However, prior works on quantifying memorization require access to the precise original data or incur substantial computational overhead, making it difficult for applications in real-world language models. To this end, we propose a fine-grained, entity-level definition to quantify memorization with conditions and metrics closer to real-world scenarios. In addition, we also present an approach for efficiently extracting sensitive entities from autoregressive language models. We conduct extensive experiments based on the proposed, probing language models' ability to reconstruct sensitive entities under different settings. We find that language models have strong memorization at the entity level and are able to reproduce the training data even with partial leakages. The results demonstrate that LLMs not only memorize their training data but also understand associations between entities. These findings necessitate that trainers of LLMs exercise greater prudence regarding model memorization, adopting memorization mitigation techniques to preclude privacy violations.

Similar Papers
  • Preprint Article

Modelling Implicit Bias in Gender–Career Associations: A systematic comparison of language models

  • May 22, 2025
  • Alexander Porshnev +5
  • Supplementary Content

Large language and vision-language models for robot: safety challenges, mitigation strategies and future directions

  • Jul 29, 2025
  • Industrial Robot: the international journal of robotics research and application
  • Xiangyu Hu +1
  • Research Article
  • Citations1

Optimization of traditional methods for determining the similarity of project names and purchases using large language models

  • Apr 01, 2024
  • Litera
  • Aleksei Aleksandrovich Golikov +2
  • Research Article
  • Citations2

Toward Cross-Hospital Deployment of Natural Language Processing Systems: Model Development and Validation of Fine-Tuned Large Language Models for Disease Name Recognition in Japanese

  • Jul 08, 2025
  • JMIR Medical Informatics
  • Seiji Shimizu +4
  • Conference Article

Data-Efficient Tabular Classification with Transformer-Based Small Language Models

  • Nov 10, 2025
  • Mario Haddad-Neto +3
  • Research Article

Research and selection of Large Learning Models for automation of ABAP-code migration

  • Sep 24, 2025
  • Management of Development of Complex Systems
  • Oleg Pozdnyakov +1
  • Research Article

Advances, Evaluation, and Explainability of Large Language Models in Healthcare: A Systematic Review

  • Dec 23, 2025
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Syed Umar Amin +2
  • Book Chapter
  • Citations3

Comparative Analysis of Air Quality Index Using Large Language Models and Machine Learning

  • Nov 28, 2024
  • Shanmugam Sundaramurthy +2
  • Conference Article
  • Citations22

The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks

  • Dec 02, 2024
  • Xiaoyi Chen +9
  • Conference Article
  • Citations18

Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

  • Jan 01, 2024
  • Daniel P Jeong +3
  • Research Article
  • Citations42

Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model.

  • Aug 21, 2024
  • PLOS digital health
  • David Soong +10
  • Research Article
  • Citations178

The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts

  • Jun 15, 2023
  • Research Ethics
  • Mohammad Hosseini +2
  • Research Article
  • Citations41

Evaluating language models for mathematics through interactions

  • Jun 03, 2024
  • Proceedings of the National Academy of Sciences
  • Katherine M Collins +13
  • Research Article
  • Citations2

Adaptive Boosting LLMs for Text Classification.

  • Jan 01, 2026
  • IEEE transactions on neural networks and learning systems
  • Mengyao Wang +6
  • PDF
  • Research Article
  • Citations26

Large language models for error detection in radiology reports: a comparative analysis between closed-source and privacy-compliant open-source models

  • Feb 20, 2025
  • European Radiology
  • Babak Salam +12
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.