- Research Article
2
- 10.1016/j.csl.2023.101569
Two evaluations on Ontology-style relation annotations
- Sep 30, 2023
- Computer Speech & Language
- Savong Bou + 2 more +2
Two evaluations on Ontology-style relation annotations
DWIE: An entity-centric dataset for multi-task document-level information extraction
Two evaluations on Ontology-style relation annotations
Two evaluations on Ontology-style relation annotations
Clinical named entity recognition and relation extraction using natural language processing of medical free text: A systematic review
BackgroundNatural Language Processing (NLP) applications have developed over the past years in various fields including its application to clinical free text for named entity recognition and relation extraction. However, there has been rapid developments the last few years that there's currently no overview of it. Moreover, it is unclear how these models and tools have been translated into clinical practice. We aim to synthesize and review these developments. MethodsWe reviewed literature from 2010 to date, searching PubMed, Scopus, the Association of Computational Linguistics (ACL), and Association of Computer Machinery (ACM) libraries for studies of NLP systems performing general-purpose (i.e., not disease- or treatment-specific) information extraction and relation extraction tasks in unstructured clinical text (e.g., discharge summaries). ResultsWe included in the review 94 studies with 30 studies published in the last three years. Machine learning methods were used in 68 studies, rule-based in 5 studies, and both in 22 studies. 63 studies focused on Named Entity Recognition, 13 on Relation Extraction and 18 performed both. The most frequently extracted entities were “problem”, “test” and “treatment”. 72 studies used public datasets and 22 studies used proprietary datasets alone. Only 14 studies defined clearly a clinical or information task to be addressed by the system and just three studies reported its use outside the experimental setting. Only 7 studies shared a pre-trained model and only 8 an available software tool. DiscussionMachine learning-based methods have dominated the NLP field on information extraction tasks. More recently, Transformer-based language models are taking the lead and showing the strongest performance. However, these developments are mostly based on a few datasets and generic annotations, with very few real-world use cases. This may raise questions about the generalizability of findings, translation into practice and highlights the need for robust clinical evaluation.
Read moreIntegrating deep learning architectures for enhanced biomedical relation extraction: a pipeline approach.
Biomedical relation extraction from scientific publications is a key task in biomedical natural language processing (NLP) and can facilitate the creation of large knowledge bases, enable more efficient knowledge discovery, and accelerate evidence synthesis. In this paper, building upon our previous effort in the BioCreative VIII BioRED Track, we propose an enhanced end-to-end pipeline approach for biomedical relation extraction (RE) and novelty detection (ND) that effectively leverages existing datasets and integrates state-of-the-art deep learning methods. Our pipeline consists of four tasks performed sequentially: named entity recognition (NER), entity linking (EL), RE, and ND. We trained models using the BioRED benchmark corpus that was the basis of the shared task. We explored several methods for each task and combinations thereof: for NER, we compared a BERT-based sequence labeling model that uses the BIO scheme with a span classification model. For EL, we trained a convolutional neural network model for diseases and chemicals and used an existing tool, PubTator 3.0, for mapping other entity types. For RE and ND, we adapted the BERT-based, sentence-bound PURE model to bidirectional and document-level extraction. We also performed extensive hyperparameter tuning to improve model performance. We obtained our best performance using BERT-based models for NER, RE, and ND, and the hybrid approach for EL. Our enhanced and optimized pipeline showed substantial improvement compared to our shared task submission, NER: 93.53 (+3.09), EL: 83.87 (+9.73), RE: 46.18 (+15.67), and ND: 38.86 (+14.9). While the performances of the NER and EL models are reasonably high, RE and ND tasks remain challenging at the document level. Further enhancements to the dataset could enable more accurate and useful models for practical use. We provide our models and code at https://github.com/janinaj/e2eBioMedRE/. Database URL: https://github.com/janinaj/e2eBioMedRE/.
Read moreSCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high training costs and difficulties in aligning with LLM preferences. To address these issues, we propose a novel universal IE paradigm—the Self-Correcting Iterative Refinement (SCIR) framework—along with a Multi-task Bilingual (Chinese-English) Self-Correcting (MBSC) dataset containing over 100,000 entries. The SCIR framework achieves plug-and-play compatibility with existing LLMs and IE systems through its Dual-Path Self-Correcting module and feedback-driven optimization, thereby significantly reducing training costs. Concurrently, the MBSC dataset tackles the challenge of preference alignment by indirectly distilling GPT-4's capabilities into IE result detection models. Experimental results demonstrate that SCIR outperforms state-of-the-art IE methods across three key tasks— named entity recognition, relation extraction, and event extraction—achieving a 5.27 percent average improvement in span-based Micro-F1 while reducing training costs by 87 percent compared to baseline approaches. These advancements not only enhance the flexibility and accuracy of IE systems but also pave the way for lightweight and efficient IE paradigms.
Read moreExtracting Reproductive Condition and Habitat Information from Text Using a Transformer-based Information Extraction Pipeline
Understanding the biology underpinning the natural regeneration of plant species in order to make plans for effective reforestation is a complex task. This can be aided by providing access to databases that contain long-term and wide-scale geographical information on species distribution, habitat, and reproduction. Although there exists widely-used biodiversity databases that contain structured information on species and their occurrences, such as the Global Biodiversity Information Facility (GBIF) and the Atlas of Living Australia (ALA), the bulk of knowledge about biodiversity still remains embedded in textual documents. Unstructured information can be made more accessible and useful for large-scale studies if there are tools and services that automatically extract meaningful information from text and store it in structured formats, e.g., open biodiversity databases, ready to be consumed for analysis (Thessen et al. 2022). We aim to enrich biodiversity occurrence databases with information on species reproductive condition and habitat, derived from text. In previous work, we developed unsupervised approaches to extract related habitats and their locations, and related reproductive condition and temporal expressions (Gabud and Batista-Navarro 2018). We built a new unsupervised hybrid approach for relation extraction (RE), which is a combination of classical rule-based pattern-matching methods and transformer-based language models that framed our RE task as a natural language inference (NLI) task. Using our hybrid approach for RE, we were able to extract related biodiversity entities from text even without a large training dataset. In this work, we implement an information extraction (IE) pipeline comprised of a named entity recognition (NER) tool and our hybrid relation extraction (RE) tool. The NER tool is a transformer-based language model that was pretrained on scientific text and then fine-tuned using COPIOUS (Conserving Philippine Biodiversity by Understanding big data; Nguyen et al. 2019), a gold standard corpus containing named entities relevant to species occurrence. We applied the NER tool to automatically annotate geographical location, temporal expression and habitat information contained within sentences. A dictionary-based approach is then used to identify mentions of reproductive conditions in text (e.g., phrases such as "fruited heavily" and "mass flowering"). We then use our hybrid RE tool to extract reproductive condition - temporal expression and habitat - geographical location entity pairs. We test our IE pipeline on the forestry compendium available in the CABI Digital Library (Centre for Agricultural and Biosciences International), and show that our work enables the enrichment of descriptive information on reproductive and habitat conditions of species. This work is a step towards enhancing a biodiversity database with the inclusion of habitat and reproductive condition information extracted from text.
Read moreMMRAG: multi-mode retrieval-augmented generation with large language models for biomedical in-context learning.
To optimize in-context learning in biomedical natural language processing by improving example selection. We introduce a novel multi-mode retrieval-augmented generation (MMRAG) framework, which integrates 4 retrieval strategies: (1) Random Mode, selecting examples arbitrarily; (2) Top Mode, retrieving the most relevant examples based on similarity; (3) Diversity Mode, ensuring variation in selected examples; and (4) Class Mode, selecting category-representative examples. This study evaluates MMRAG on 3 core biomedical NLP tasks: Named Entity Recognition (NER), Relation Extraction (RE), and Text Classification (TC). The datasets used include BC2GM for gene and protein mention recognition (NER), DDI for drug-drug interaction extraction (RE), GIT for general biomedical information extraction (RE), and HealthAdvice for health-related text classification (TC). The framework is tested with 2 large language models (Llama-2-7B and Llama-3-8B) and 3 retrievers (Contriever, MedCPT, and BGE-Large) to assess performance across different retrieval strategies. The results from the Random Mode indicate that providing more examples in the prompt improves the model's generation performance. Meanwhile, Top Mode and Diversity Mode significantly outperform Random Mode on the RE (DDI) task, achieving an F1 score of 0.9669-a 26.4% improvement. Among the 3 retrievers tested, Contriever outperformed the other 2 in a greater number of experiments. Additionally, Llama 2 and Llama 3 demonstrated varying capabilities across different tasks, with Llama 3 showing a clear advantage in handling NER tasks. MMRAG effectively enhances biomedical in-context learning by refining example selection, mitigating data scarcity issues, and demonstrating superior adaptability for NLP-driven healthcare applications.
Read moreInformation Extraction with Negative Examples for Author Biographies in Scientific Literatures
Information extraction (IE) is a textual information processing task concerned with the automatic extraction of entity mentions and relational structures from documents, which has been widely studied and applied in various fields such as biomedicine and news. In this paper, we extend this technique to biographies and propose a new BioIE3 dataset to extract the authors' experiences (including birth, education, work and awards) from their profiles in the scientific literatures. Named entity recognition (NER) and relation extraction (RE) tasks are performed simultaneously in a neural architecture based on the Skip-Gram Word2vec, Bidirectional Long Short Term Memory (BiLSTM) and Conditional Random Fields (CRF) models. On the basis of the above work, negative examples are especially introduced into the training stage to effectively improve and increase the robustness of the model.
Read moreA region-based hypergraph network for joint entity-relation extraction
A region-based hypergraph network for joint entity-relation extraction
Domain-Specific Knowledge Graph for Quality Engineering of Continuous Casting: Joint Extraction-Based Construction and Adversarial Training Enhanced Alignment
The intelligent development of continuous casting quality engineering is an essential step for the efficient production of high-quality billets. However, there are many quality defects that require strong expertise for handling. In order to reduce reliance on expert experience and improve the intelligent management level of billet quality knowledge, we focus on constructing a Domain-Specific Knowledge Graph (DSKG) for the quality engineering of continuous casting. To achieve joint extraction of billet quality defects entity and relation, we propose a Self-Attention Partition and Recombination Model (SAPRM). SAPRM divides domain-specific sentences into three parts: entity-related, relation-related, and shared features, which are specifically for Named Entity Recognition (NER) and Relation Extraction (RE) tasks. Furthermore, for issues of entity ambiguity and repetition in triples, we propose a semi-supervised incremental learning method for knowledge alignment, where we leverage adversarial training to enhance the performance of knowledge alignment. In the experiment, in the knowledge extraction part, the NER and RE precision of our model achieved 86.7% and 79.48%, respectively. RE precision improved by 20.83% compared to the baseline with sequence labeling method. Additionally, in the knowledge alignment part, the precision of our model reached 99.29%, representing a 1.42% improvement over baseline methods. Consequently, the proposed model with the partition mechanism can effectively extract domain knowledge, cand the semi-supervised method can take advantage of unlabeled triples. Our method can adapt the domain features and construct a high-quality knowledge graph for the quality engineering of continuous casting, providing an efficient solution for billet defect issues.
Read moreProgressive Multi-task Learning with Controlled Information Flow for Joint Entity and Relation Extraction
Multitask learning has shown promising performance in learning multiple related tasks simultaneously, and variants of model architectures have been proposed, especially for supervised classification problems. One goal of multitask learning is to extract a good representation that sufficiently captures the relevant part of the input about the output for each learning task. To achieve this objective, in this paper we design a multitask learning architecture based on the observation that correlations exist between outputs of some related tasks (e.g. entity recognition and relation extraction tasks), and they reflect the relevant features that need to be extracted from the input. As outputs are unobserved, our proposed model exploits task predictions in lower layers of the neural model, also referred to as early predictions in this work. But we control the injection of early predictions to ensure that we extract good task-specific representations for classification. We refer to this model as a Progressive Multitask learning model with Explicit Interactions (PMEI). Extensive experiments on multiple benchmark datasets produce state-of-the-art results on the joint entity and relation extraction task.
Read moreSmart Contracts Auto-generation for Supply Chain Contexts
The introduction of blockchain technology into Supply Chain management has opened the possibility of faster and more secure transactions of commodities and services. As for every blockchain, Smart Contracts are the tool for controlling the transactions in blockchain-based supply chains. In this paper, we introduce a method for automating the implementation of natural language contracts into Smart Contracts in the Supply Chain context. The basic idea here is to extract information from a natural language contract using two Natural Language Processing (NLP) techniques, the Named Entity Recognition (NER) and Relation Extraction (RE), and then use this extracted information to automatically create a corresponding Smart Contract. This is an ongoing project, and we implemented the first phase of NLP, i.e., NER. The main issue we are facing here is the limited availability of annotated contract datasets. To tackle this challenge, we created an annotated legal contract dataset dedicated to the NER task. The dataset is analyzed with the deep learning method (BiLSTM) and transformer-based method (BERT). As per the generation of smart contracts, our approach consists of identifying meaningful entities and the relations between them and then representing them as business logic that can be directly incorporated into computer code as blockchain smart contracts.KeywordsNLPNERREDatasetDeep learningLegal domain
Read moreJoint entity and relation extraction based on a hybrid neural network
Joint entity and relation extraction based on a hybrid neural network
Influence of Knowledge Extraction Methods on the Effectiveness of Graph-Based Rag Systems
This paper investigates the impact of knowledge extraction methods on the effectiveness of RAG (Retrieval-Augmented Generation) systems that utilize knowledge graphs. It highlights that the quality of the knowledge graph, which is formed using various knowledge extraction methods, is crucial for overcoming limitations of large language models (LLMs), such as “hallucinations”. The paper analyzes the architectures of LightRAG and GraphRAG, emphasizing that the selection of an optimal knowledge extraction strategy depends on specific tasks and the subject area. LLMs have advanced significantly, but they have limitations, including generating factually incorrect information (“hallucinations”) and possessing “outdated knowledge”. RAG systems were proposed to address these issues by combining LLMs with external knowledge bases. This approach reduces hallucinations, ensures factual accuracy, solves the problem of outdated knowledge, and increases transparency. Knowledge graphs are powerful tools for structuring information, consisting of entities (nodes) and relations (edges). They enhance RAG systems by enabling more precise and contextually grounded retrieval compared to keyword-based searches. The quality of a knowledge graph depends on the knowledge extraction methods used, which include named entity recognition (NER), relation extraction (RE), entity linking (EL), and event extraction. Different methods, such as rule-based, classical machine learning, and deep learning approaches, have varying trade-offs in terms of accuracy, completeness, and scalability. Entity linking and knowledge graph completion are also crucial for accuracy and richness. LightRAG and GraphRAG are two main graph-based RAG systems. LightRAG uses the knowledge graph as a quick reference, requiring high precision in knowledge extraction to avoid noise degradation. GraphRAG uses the knowledge graph as a domain model, where completeness of extracted knowledge is more critical, though systematic errors are still harmful. Both systems rely on LLMs for knowledge extraction, which makes them dependent on the LLM’s quality and the size of document fragments processed. The theoretical analysis confirms that the effectiveness of RAG systems is critically dependent on knowledge extraction methods. The quality, completeness, and accuracy of the knowledge graph directly influence the RAG system’s ability to provide relevant, accurate, and truthful answers. Different RAG system architectures like LightRAG and GraphRAG have distinct requirements for knowledge graph characteristics, prioritizing either accuracy or completeness in knowledge extraction.
Read moreJoint Extraction Model of Entity Relations Based on BERT-CRF
Entity and relation extraction (ERE) is a major task in information extraction and knowledge mapping. Existing methods usually consider two tasks, named entity recognition (NER) and relation extraction (RE), separately using a pipeline approach, which loses a lot of interaction information between tasks and contextual information of text sequences. In order to settle this problem, this paper proposes an end-to-end entity relation joint extraction method based on the head-entity attention mechanism and fusing contextual semantic features. The overall structure of this method adopts BERT-CRF to decode the header entity and its type, and then uses the header entity information as the Query in the attention mechanism, while fusing entity type label embedding and entity relative position to achieve feature enhancement, which enhances the information interaction between entity model and relational model. In the experiments of the commonly used English dataset NYT and Chinese dataset DuIE, this method has achieved high extraction accuracy and F1 score. It is shown that the model is applicable in both English and Chinese contexts.
Read moreMulti-View Consistency for Relation Extraction via Mutual Information and Structure Prediction
Relation Extraction (RE) is one of the fundamental tasks in Information Extraction. The goal of this task is to find the semantic relations between entity mentions in text. It has been shown in many previous work that the structure of the sentences (i.e., dependency trees) can provide important information/features for the RE models. However, the common limitation of the previous work on RE is the reliance on some external parsers to obtain the syntactic trees for the sentence structures. On the one hand, it is not guaranteed that the independent external parsers can offer the optimal sentence structures for RE and the customized structures for RE might help to further improve the performance. On the other hand, the quality of the external parsers might suffer when applied to different domains, thus also affecting the performance of the RE models on such domains. In order to overcome this issue, we introduce a novel method for RE that simultaneously induces the structures and predicts the relations for the input sentences, thus avoiding the external parsers and potentially leading to better sentence structures for RE. Our general strategy to learn the RE-specific structures is to apply two different methods to infer the structures for the input sentences (i.e., two views). We then introduce several mechanisms to encourage the structure and semantic consistencies between these two views so the effective structure and semantic representations for RE can emerge. We perform extensive experiments on the ACE 2005 and SemEval 2010 datasets to demonstrate the advantages of the proposed method, leading to the state-of-the-art performance on such datasets.
Read more