- Research Article
19
- 10.1016/j.nlp.2024.100099
A novel prompting method for few-shot NER via LLMs
- Aug 24, 2024
- Natural Language Processing Journal
- Qi Cheng + 5 more +5
A novel prompting method for few-shot NER via LLMs
In recent years, the field of natural language processing has witnessed remarkable advancements due to the success of large language models. These models leverage the Transformer architecture and pre-training techniques to achieve impressive results. In this paper, we draw inspiration from large language models and apply these techniques into the task of named entity recognition in the domain of power grids, which is critical for building power grid knowledge graphs and question-answering systems. Specifically, we propose a BERT-CNN-BIGRU-CRF deep learning model for named entity recognition. This model effectively harnesses the semantic modeling capabilities and pre-training knowledge of BERT, which is based on the Transformer architecture. By incorporating CNN and BIGRU, the model captures and models both local and global features, respectively. The CRF layer is employed for label classification. This combination of components ensures a high level of recognition accuracy. To evaluate the performance of the proposed model, we train our model on annotated maintenance plan data. We compare its results with those of other commonly used models. The evaluation metrics include recall, precision, and F1 score, which are widely employed in named entity recognition tasks. Our proposed model achieves optimal performance across all three metrics, demonstrating its superiority over other models.
A novel prompting method for few-shot NER via LLMs
A novel prompting method for few-shot NER via LLMs
#2924 Comparison of large language models and traditional natural language processing techniques in predicting arteriovenous fistula failure
Background and Aims Large language models (LLMs) have gained significant attention in the field of natural language processing (NLP), marking a shift from traditional techniques like Term Frequency-Inverse Document Frequency (TF-IDF). We developed a traditional NLP model to predict arteriovenous fistula (AVF) failure within next 30 days using clinical notes. The goal of this analysis was to investigate whether LLMs would outperform traditional NLP techniques, specifically in the context of predicting AVF failure within the next 30 days using clinical notes. Method We defined AVF failure as the change in status from active to permanently unusable status or temporarily unusable status. We used data from a large kidney care network from January 2021 to December 2021. Two models were created using LLMs and traditional TF-IDF technique. We used “distilbert-base-uncased”, a distilled version of BERT base model [1], and compared its performance with traditional TF-IDF-based NLP techniques. The dataset was randomly divided into 60% training, 20% validation and 20% test dataset. The test data, comprising of unseen patients’ data was used to evaluate the performance of the model. Both models were evaluated using metrics such as area under the receiver operating curve (AUROC), accuracy, sensitivity, and specificity. Results The incidence of 30 days AVF failure rate was 2.3% in the population. Both LLMs and traditional showed similar overall performance as summarized in Table 1. Notably, LLMs showed marginally better performance in certain evaluation metrics. Both models had same AUROC of 0.64 on test data. The accuracy and balanced accuracy for LLMs were 72.9% and 59.7%, respectively, compared to 70.9% and 59.6% for the traditional TF-IDF approach. In terms of specificity, LLMs scored 73.2%, slightly higher than the 71.2% observed for traditional NLP methods. However, LLMs had a lower sensitivity of 46.1% compared to 48% for traditional NLP. However, it is worth noting that training on LLMs took considerably longer than TF-IDF. Moreover, it also used higher computational resources such as utilization of graphics processing units (GPU) instances in cloud-based services, leading to higher cost. Conclusion In our study, we discovered that advanced LLMs perform comparably to traditional TF-IDF modeling techniques in predicting the failure of AVF. Both models demonstrated identical AUROC. While specificity was higher in LLMs compared to traditional NLP, sensitivity was higher in traditional NLP compared to LLMs. LLM was fine-tuned with a limited dataset, which could have influenced its performance to be similar to that of traditional NLP methods. This finding suggests that while LLMs may excel in certain scenarios, such as performing in-depth sentiment analysis of patient data for complex tasks, their effectiveness is highly dependent on the specific use case. It is crucial to weigh the benefits against the resources required for LLMs, as they can be significantly more resource-intensive and costly compared to traditional TF-IDF methods. This highlights the importance of a use-case-driven approach in selecting the appropriate NLP technique for healthcare applications.
Read moreAdvancements and Applications of Large Language Models in Natural Language Processing: A Comprehensive Review
Abstract. Large language models (LLMs) have revolutionized the field of natural language processing (NLP), demonstrating remarkable capabilities in understanding, generating, and manipulating human language. This comprehensive review explores the development, applications, optimizations, and challenges of LLMs. This paper begin by tracing the evolution of these models and their foundational architectures, such as the Transformer, GPT, and BERT. We then delve into the applications of LLMs in natural language understanding tasks, including sentiment analysis, named entity recognition, question answering, and text summarization, highlighting real-world use cases. Next, we examine the role of LLMs in natural language generation, covering areas such as content creation, language translation, personalized recommendations, and automated responses. We further discuss LLM applications in other NLP tasks like text style transfer, text correction, and language model pre-training. Subsequently, we explore techniques for optimizing and improving LLMs, including model compression, explainability, robustness, and security. Finally, we address the challenges posed by the significant computational requirements, sample inefficiency, and ethical considerations surrounding LLMs. We conclude by discussing potential future research directions, such as efficient architectures, few-shot learning, bias mitigation, and privacy-preserving techniques, which will shape the ongoing development and responsible deployment of LLMs in NLP.
Read moreInvestigating the role of Named Entity Recognition in Question Answering Models
Machine Reading Comprehension (MRC) is a challenging Question - Answering (QA) task that helps the user in providing the answer to the given question. There is a lot of progress in this area due to the availability of large datasets and large pre-trained language models based on transformer architecture (BERT). Named Entity Recognition (NER) was used for neural QA systems to improve performance. However, whether NER plays a vital role in a QA system built using contextual embeddings obtained through BERT variants is not explored. To fill this gap, we investigate whether NER is helpful in improving the performance of QA systems built using BERT variants. We experimented with Squad 2.0 using SpanBERT. The Squad 2.0 dataset has both answerable and unanswerable questions. The proposed model finds the answer span if the question is answerable and, provides justification for the unanswerable questions. We perform question analysis to find the expected answer tag and then use that information to find the relevant parts of the passage in order to retrieve the answer span.
Read moreMulti-Perspective Knowledge Distillation of LLM for NER in IPE Courses
Named Entity Recognition (NER) is essential for extracting meaningful entities from text, but existing methods struggle with complex linguistic structures and domain-specific contexts, such as those in Ideological and Political Education (IPE) texts. This paper proposes a novel approach using multi-perspective knowledge distillation from Large Language Models (LLMs) to enhance NER performance in IPE. The method involves constructing a specialized dataset for IPE and generating intermediate reasoning data using the Qwen14B model through a Chain-of-Thought (CoT) approach. Knowledge from the LLM is then distilled into a smaller NER model using techniques like DoRA fine-tuning and multi-perspective alignment, which includes feature, content, and distribution alignment. Experiments show significant improvements over state-of-the-art models, with gains of 3.46% in precision, 5.79% in recall, and 2.54% in F1 score. The method also demonstrates strong few-shot learning capabilities, achieving high performance with limited training data.
Read moreComparative Analysis of Large Language Models in Chinese Medical Named Entity Recognition
The emergence of large language models (LLMs) has provided robust support for application tasks across various domains, such as name entity recognition (NER) in the general domain. However, due to the particularity of the medical domain, the research on understanding and improving the effectiveness of LLMs on biomedical named entity recognition (BNER) tasks remains relatively limited, especially in the context of Chinese text. In this study, we extensively evaluate several typical LLMs, including ChatGLM2-6B, GLM-130B, GPT-3.5, and GPT-4, on the Chinese BNER task by leveraging a real-world Chinese electronic medical record (EMR) dataset and a public dataset. The experimental results demonstrate the promising yet limited performance of LLMs with zero-shot and few-shot prompt designs for Chinese BNER tasks. More importantly, instruction fine-tuning significantly enhances the performance of LLMs. The fine-tuned offline ChatGLM2-6B surpassed the performance of the task-specific model BiLSTM+CRF (BC) on the real-world dataset. The best fine-tuned model, GPT-3.5, outperforms all other LLMs on the publicly available CCKS2017 dataset, even surpassing half of the baselines; however, it still remains challenging for it to surpass the state-of-the-art task-specific models, i.e., Dictionary-guided Attention Network (DGAN). To our knowledge, this study is the first attempt to evaluate the performance of LLMs on Chinese BNER tasks, which emphasizes the prospective and transformative implications of utilizing LLMs on Chinese BNER tasks. Furthermore, we summarize our findings into a set of actionable guidelines for future researchers on how to effectively leverage LLMs to become experts in specific tasks.
Read morePerformance and Reproducibility of Large Language Models in Named Entity Recognition: Considerations for the Use in Controlled Environments
IntroductionRecent artificial intelligence (AI) advances can generate human-like responses to a wide range of queries, making them a useful tool for healthcare applications. Therefore, the potential use of large language models (LLMs) in controlled environments regarding efficacy, reproducibility, and operability will be of paramount interest.ObjectiveWe investigated if and how GPT 3.5 and GPT 4 models can be directly used as a part of a GxP validated system and compared the performance of externally hosted GPT 3.5 and GPT 4 against LLMs, which can be hosted internally. We explored zero-shot LLM performance for named entity recognition (NER) and relation extraction tasks, investigated which LLM has the best zero-shot performance to be used potentially for generating training data proposals, evaluated the LLM performance of seven entities for medical NER in zero-shot experiments, selected one model for further performance improvement (few-shot and fine-tuning: Zephyr-7b-beta), and investigated how smaller open-source LLMs perform in contrast to GPT models and to a small fine-tuned T5 Base.MethodsWe performed reproducibility experiments to evaluate if LLMs can be used in controlled environments and utilized guided generation to use the same prompt across multiple models. Few-shot learning and quantized low rank adapter (QLoRA) fine-tuning were applied to further improve LLM performance.Results and ConclusionWe demonstrated that zero-shot GPT 4 performance is comparable with a fine-tuned T5, and Zephyr performed better than zero-shot GPT 3.5, but the recognition of product combinations such as product event combination was significantly better by using a fine-tuned T5. Although Open AI launched recently GPT versions to improve the generation of consistent output, both GPT variants failed to demonstrate reproducible results. The lack of reproducibility together with limitations of external hosted systems to keep validated systems in a state of control may affect the use of closed and proprietary models in regulated environments. However, due to the good NER performance, we recommend using GPT for creating annotation proposals for training data as a basis for fine-tuning.
Read moreFsPONER: Few-Shot Prompt Optimization for Named Entity Recognition in Domain-Specific Scenarios
Large Language Models (LLMs) have provided a new pathway for Named Entity Recognition (NER) tasks. Compared with fine-tuning, LLM-powered prompting methods avoid the need for training, conserve substantial computational resources, and rely on minimal annotated data. Previous studies have achieved comparable performance to fully supervised BERT-based fine-tuning approaches on general NER benchmarks. However, none of the previous approaches has investigated the efficiency of LLM-based few-shot learning in domain-specific scenarios. To address this gap, we introduce FsPONER, a novel approach for optimizing few-shot prompts, and evaluate its performance on domain-specific NER datasets, with a focus on industrial manufacturing and maintenance, while using multiple LLMs – GPT-4-32K, GPT-3.5-Turbo, LLaMA 2-chat, and Vicuna. FsPONER consists of three few-shot selection methods based on random sampling, TF-IDF vectors, and a combination of both. We compare these methods with a general-purpose GPT-NER method as the number of few-shot examples increases and evaluate their optimal NER performance against fine-tuned BERT and LLaMA 2-chat. In the considered real-world scenarios with data scarcity, FsPONER with TF-IDF surpasses fine-tuned models by approximately 10% in F1 score.
Read moreOn the role of the UMLS in supporting diagnosis generation proposed by Large Language Models
On the role of the UMLS in supporting diagnosis generation proposed by Large Language Models
Teaching Machines to Find Names
In the field of Natural Language Processing, one of the very important research areas of Information Extraction (IE) comes in Named Entity Recognition (NER). NER is a subtask of IE that seeks to identify and classify the predefined categories of named entities in text documents. Considerable amount of work has been done on NER in recent years due to the increasing demand of automated texts and the wide availability of electronic corpora. While it is relatively easy and natural for a human reader to read and understand the context of a given article, getting a machine to understand and differentiate between words is a big challenge. For instance, the word ‘brown’ may refer to a person called Mr. Brown, or the colour of an item which is brown. Human readers can easily discern the meaning of the word by looking at the context of that particular sentence, but it would be almost impossible for a computer to interpret it without any additional information. To deal with the issue, researchers in NER field have proposed various rule-based systems (Wakao, Gaizauskas & Wilks, 1996; Krupka & Hausman, 1998; Maynard, Tablan, Ursu, Cunningham & Wilks, 2001). These systems are able to achieve high accuracy in recognition with the help of some lists of known named entities called gazetteers. The problem with rule-based approach is that it lacks the robustness and portability. It incurs steep maintenance cost especially when new rules need to be introduced for some new information or new domains. A better option is thus to use machine learning approach that is trainable and adaptable. Three wellknown machine learning approaches that have been used extensively in NER are Hidden Markov Model (HMM), Maximum Entropy Model (MEM) and Decision Tree. Many of the existing machine learning-based NER systems (Bikel, Schwartz & Weischedel, 1999; Zhou & Su, 2002; Borthwick, Sterling, Agichten & Grisham, 1998; Bender, Och & Ney, 2003; Chieu & Ng, 2002; Sekine, Grisham & Shinnou, 1998) are able to achieve near-human performance for named entity tagging, even though the overall performance is still about 2% short from the rule-based systems. There have also been many attempts to improve the performance of NER using a hybrid approach with the combination of handcrafted rules and statistical models (Mikheev, Moens & Grover, 1999; Srihari & Li, 2000; Seon, Ko, Kim & Seo, 2001). These systems can achieve relatively good performance in the targeted domains owing to the comprehensive handcrafted rules. Nevertheless, the portability problem still remains unsolved when it comes to dealing with NER in various domains. As such, this article presents a hybrid machine learning approach using MEM and HMM successively. The reason for using two statistical models in succession instead of one is due to the distinctive nature of the two models. HMM is able to achieve better performance than any other statistical models, and is generally regarded as the most successful one in machine learning approach. However, it suffers from sparseness problem, which means considerable amount of data is needed for it to achieve acceptable performance. On the other hand, MEM is able to maintain reasonable performance even when there is little data available for training purpose. The idea is therefore to walkthrough the testing corpus using MEM first in order to generate a temporary tagging result, while this procedure can be simultaneously used as a training process for HMM. During the second walkthrough, the corpus uses HMM for the final tagging. In this process, the temporary tagging result generated by MEM will be used as a reference for subsequent error checking and correction. In the case when there is little training data available, the final result can still be reliable based on the contribution of the initial MEM tagging result.
Read moreROFED-LLM: Robust Federated Learning for Large Language Models in Adversarial Wireless Environments
Large language models (LLMs) have made significant advances in the field of natural language processing (NLP). However, their centralized training approach faces challenges related to data privacy, communication efficiency, and robustness against adversarial attacks, particularly in wireless environments. With the gradual depletion of high-quality public data, there is an urgent need to leverage private data distributed across various parties. Although federated learning (FL) offers a privacy-preserving collaborative training paradigm, it struggles to meet the high computational demands of edge devices and remains vulnerable to adversarial attacks. This paper introduces <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ROFED</small>-LLM, a novel framework for robust, privacy-preserving training of LLMs on decentralized private data over wireless networks. By integrating split federated learning, which partitions the model across devices to enhance privacy, with adaptive jamming defense mechanisms, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ROFED</small>-LLM enables collaborative LLM training without raw data sharing while ensuring resilience against wireless adversarial attacks. Our multi-modal defense strategy combines model-level protections, such as differential privacy and dynamic pruning, with communication-level safeguards, including adaptive beamforming which optimizes wireless signal transmission to mitigate interference, and resource allocation optimization. Extensive experiments across diverse NLP tasks demonstrate <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ROFED</small>-LLM's superiority, achieving a 12.87% improvement in privacy preservation and 18.26% enhancement in jamming resilience compared to existing methods such as FedAvg and SCAFFOLD, with only a marginal 3.94% trade-off in model accuracy. Our code repository has been open sourced at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://anonymous.4open.science/r/RoFed-LLM-54E1</uri>.
Read moreCharacter-Boundary-Aware and Type-Decoupled Named Entity Recognition Enhanced by a Post-Correction Model
Named Entity Recognition (NER) aims to identify entities with specific semantic meanings from text and classify them into predefined categories. With the rapid progress of generative large language models (LLMs), their strong capabilities in text comprehension and generation have sparked growing interest in applying them to information extraction tasks, including NER, within a generative paradigm. However, LLMs are often regarded as ”black-box” systems, making it difficult to interpret the reasoning behind their predictions in entity recognition. In contrast, traditional BERT-based models explicitly assign labels to each token, ensuring precise capture of entity boundaries, while LLMs tend to infer spans from context, which may weaken their sensitivity to boundary information.To overcome these challenges, we propose a novel framework: CBATD-PC. Our approach integrates an instruction-tuning strategy for model construction and introduces a character-level boundary labeling mechanism. By decoupling the NER process into two stages, boundary detection and type classification, we establish a more structured and interpretable prediction pipeline. Moreover, by leveraging the contextual reasoning strengths of LLMs, we design a post-correction module that fine-tunes and refines the initial extraction results, thereby improving the overall accuracy and robustness of entity recognition.
Read moreMitigating Hallucination in Large Language Models: Techniques, Applications, and Implications
Large Language Models (LLMs) such as GPT, LLaMA, and PaLM have transformed the field of Natural Language Processing (NLP) by achieving remarkable results in text generation, summarization, translation, question answering, and dialogue systems. Their wide adoption across industries highlights their usefulness but also exposes a critical limitation—hallucination. Hallucination occurs when models generate information that is false, misleading, or fabricated. These errors can vary from small factual mistakes, like incorrect dates or figures, to serious inaccuracies that may cause harm in sensitive areas such as healthcare, education, and software development. This paper explores the concept and classification of hallucinations in LLMs, examines techniques to reduce them—including prompt engineering, fine-tuning, and Retrieval-Augmented Generation (RAG)—and discusses ethical implications and real-world applications. By comparing multiple strategies, the study aims to contribute to developing more reliable and trustworthy AI systems.
Read moreSmall Language Model Makes an Effective Long Text Extractor
Named Entity Recognition (NER) is a fundamental problem in natural language processing (NLP). However, the task of extracting longer entity spans (e.g., awards) from extended texts (e.g., homepages) is barely explored. Current NER methods predominantly fall into two categories: span-based methods and generation-based methods. Span-based methods require the enumeration of all possible token-pair spans, followed by classification on each span, resulting in substantial redundant computations and excessive GPU memory usage. In contrast, generation-based methods involve prompting or fine-tuning large language models (LLMs) to adapt to downstream NER tasks. However, these methods struggle with the accurate generation of longer spans and often incur significant time costs for effective finetuning. To address these challenges, this paper introduces a lightweight span-based NER method called SeNER, which incorporates a bidirectional arrow attention mechanism coupled with LogN-Scaling on the [CLS] token to embed long texts effectively, and comprises a novel bidirectional sliding-window plus-shaped attention (BiSPA) mechanism to reduce redundant candidate token-pair spans significantly and model interactions between token-pair spans simultaneously. Extensive experiments demonstrate that our method achieves state-of-the-art extraction accuracy on three long NER datasets and is capable of extracting entities from long texts in a GPU-memory-friendly manner.
Read moreUsing Synthetic Health Care Data to Leverage Large Language Models for Named Entity Recognition: Development and Validation Study.
Named entity recognition (NER) plays a vital role in extracting critical medical entities from health care records, facilitating applications such as clinical decision support and data mining. Developing robust NER models for low-resource languages, such as Estonian, remains a challenge due to the scarcity of annotated data and domain-specific pretrained models. Large language models (LLMs) have proven to be promising in understanding text from any language or domain. This study addresses the development of medical NER models for low-resource languages, specifically Estonian. We propose a novel approach by generating synthetic health care data and using LLMs to annotate them. These synthetic data are then used to train a high-performing NER model, which is applied to real-world medical texts, preserving patient data privacy. Our approach to overcoming the shortage of annotated Estonian health care texts involves a three-step pipeline: (1) synthetic health care data are generated using a locally trained GPT-2 model on Estonian medical records, (2) the synthetic data are annotated with LLMs, specifically GPT-3.5-Turbo and GPT-4, and (3) the annotated synthetic data are then used to fine-tune an NER model, which is later tested on real-world medical data. This paper compares the performance of different prompts; assesses the impact of GPT-3.5-Turbo, GPT-4, and a local LLM; and explores the relationship between the amount of annotated synthetic data and model performance. The proposed methodology demonstrates significant potential in extracting named entities from real-world medical texts. Our top-performing setup achieved an F1-score of 0.69 for drug extraction and 0.38 for procedure extraction. These results indicate a strong performance in recognizing certain entity types while highlighting the complexity of extracting procedures. This paper demonstrates a successful approach to leveraging LLMs for training NER models using synthetic data, effectively preserving patient privacy. By avoiding reliance on human-annotated data, our method shows promise in developing models for low-resource languages, such as Estonian. Future work will focus on refining the synthetic data generation and expanding the method's applicability to other domains and languages.
Read more