• Home
  • Search
  • Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation
  • Open Access IconOpen Access
  • Cite Icon4
  • https://doi.org/10.18653/v1/2024.acl-long.124Copy DOI Icon

Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation

  • Jan 1, 2024
  • Xunjian Yin +3 more
Show More
  • Abstract
  • Literature Map
  • Citations
  • Similar Papers
Abstract

In recent years, substantial advancements have been made in the development of large language models, achieving remarkable performance across diverse tasks. To evaluate the knowledge ability of language models, previous studies have proposed lots of benchmarks based on question-answering pairs. We argue that it is not reliable and comprehensive to evaluate language models with a fixed question or limited paraphrases as the query, since language models are sensitive to prompt. Therefore, we introduce a novel concept named knowledge boundary to encompass both prompt-agnostic and promptsensitive knowledge within language models. Knowledge boundary avoids prompt sensitivity in language model evaluations, rendering them more dependable and robust. To explore the knowledge boundary for a given model, we propose a projected gradient descent method with semantic constraints, a new algorithm designed to identify the optimal prompt for each piece of knowledge. Experiments demonstrate a superior performance of our algorithm in computing the knowledge boundary compared to existing methods. Furthermore, we evaluate the ability of multiple language models in several domains with knowledge boundary.

Similar Papers
  • Preprint Article

Modelling Implicit Bias in Gender–Career Associations: A systematic comparison of language models

  • May 22, 2025
  • Alexander Porshnev +5
  • Supplementary Content

Large language and vision-language models for robot: safety challenges, mitigation strategies and future directions

  • Jul 29, 2025
  • Industrial Robot: the international journal of robotics research and application
  • Xiangyu Hu +1
  • Research Article
  • Citations1

Optimization of traditional methods for determining the similarity of project names and purchases using large language models

  • Apr 01, 2024
  • Litera
  • Aleksei Aleksandrovich Golikov +2
  • Conference Article

Data-Efficient Tabular Classification with Transformer-Based Small Language Models

  • Nov 10, 2025
  • Mario Haddad-Neto +3
  • Research Article

Research and selection of Large Learning Models for automation of ABAP-code migration

  • Sep 24, 2025
  • Management of Development of Complex Systems
  • Oleg Pozdnyakov +1
  • Research Article

Advances, Evaluation, and Explainability of Large Language Models in Healthcare: A Systematic Review

  • Dec 23, 2025
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Syed Umar Amin +2
  • Research Article
  • Citations4

Quantifying and Analyzing Entity-Level Memorization in Large Language Models

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Zhenhong Zhou +3
  • Research Article
  • Citations41

Evaluating language models for mathematics through interactions

  • Jun 03, 2024
  • Proceedings of the National Academy of Sciences
  • Katherine M Collins +13
  • Conference Article

KnowFC: Navigating Knowledge Conflicts in Large Language Model-based Fact-Checking

  • Feb 16, 2026
  • Yue Zhang +8
  • Conference Article
  • Citations18

Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

  • Jan 01, 2024
  • Daniel P Jeong +3
  • Research Article
  • Citations42

Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model.

  • Aug 21, 2024
  • PLOS digital health
  • David Soong +10
  • Research Article
  • Citations178

The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts

  • Jun 15, 2023
  • Research Ethics
  • Mohammad Hosseini +2
  • Conference Article
  • Citations196

Structured Pruning of Large Language Models

  • Jan 01, 2020
  • Ziheng Wang +2
  • Research Article
  • Citations2

Adaptive Boosting LLMs for Text Classification.

  • Jan 01, 2026
  • IEEE transactions on neural networks and learning systems
  • Mengyao Wang +6
  • PDF
  • Research Article
  • Citations26

Large language models for error detection in radiology reports: a comparative analysis between closed-source and privacy-compliant open-source models

  • Feb 20, 2025
  • European Radiology
  • Babak Salam +12
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.