• Home
  • Search
  • SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
  • https://doi.org/10.1609/aaai.v40i36.40326Copy DOI Icon

SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema

  • Mar 14, 2026
  • Yu Fang +5 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high training costs and difficulties in aligning with LLM preferences. To address these issues, we propose a novel universal IE paradigm—the Self-Correcting Iterative Refinement (SCIR) framework—along with a Multi-task Bilingual (Chinese-English) Self-Correcting (MBSC) dataset containing over 100,000 entries. The SCIR framework achieves plug-and-play compatibility with existing LLMs and IE systems through its Dual-Path Self-Correcting module and feedback-driven optimization, thereby significantly reducing training costs. Concurrently, the MBSC dataset tackles the challenge of preference alignment by indirectly distilling GPT-4's capabilities into IE result detection models. Experimental results demonstrate that SCIR outperforms state-of-the-art IE methods across three key tasks— named entity recognition, relation extraction, and event extraction—achieving a 5.27 percent average improvement in span-based Micro-F1 while reducing training costs by 87 percent compared to baseline approaches. These advancements not only enhance the flexibility and accuracy of IE systems but also pave the way for lightweight and efficient IE paradigms.

Similar Papers
  • Conference Article
  • Citations49

Inducing information extraction systems for new languages via cross-language projection

  • Jan 01, 2002
  • Ellen Riloff +2
  • Research Article
  • Citations64

InfoXtract: A customizable intermediate level information extraction engine

  • Jan 01, 2008
  • Natural Language Engineering
  • Rohini K Srihari +3
  • Video Transcripts

Cost-effective End-to-end Information Extraction for Semi-structured Document Images

  • Oct 15, 2021
  • Underline Science Inc.
  • Hyunji Lee +4
  • Research Article
  • Citations64

DWIE: An entity-centric dataset for multi-task document-level information extraction

  • Mar 20, 2021
  • Information Processing & Management
  • Klim Zaporojets +3
  • Research Article
  • Citations7

MMRAG: multi-mode retrieval-augmented generation with large language models for biomedical in-context learning.

  • Aug 04, 2025
  • Journal of the American Medical Informatics Association : JAMIA
  • Zaifu Zhan +4
  • PDF
  • Research Article
  • Citations64

BIOSMILE: a semantic role labeling system for biomedical verbs using a maximum-entropy model with automatically generated template features.

  • Sep 01, 2007
  • BMC Bioinformatics
  • Richard Tzong-Han Tsai +9
  • Research Article

Influence of Knowledge Extraction Methods on the Effectiveness of Graph-Based Rag Systems

  • Nov 26, 2025
  • NaUKMA Research Papers. Computer Science
  • Maksym Androshchuk
  • Research Article
  • Citations7

Performance and Reproducibility of Large Language Models in Named Entity Recognition: Considerations for the Use in Controlled Environments

  • Dec 11, 2024
  • Drug Safety
  • Jürgen Dietrich +1
  • Preprint Article
  • Citations1

Utilizing Large Language Models for Geoscience Literature Information Extraction

  • Jan 20, 2025
  • Peng Yu +3
  • Research Article
  • Citations68

Clinical named entity recognition and relation extraction using natural language processing of medical free text: A systematic review

  • Jun 05, 2023
  • International journal of medical informatics
  • David Fraile Navarro +6
  • Conference Article
  • Citations4

A Novel Method of Chinese Web Information Extraction and Applications

  • Jul 01, 2009
  • Zhong Liu +1
  • Research Article
  • Citations14

Information Extraction Strategies for Thai Documents

  • Jun 01, 2001
  • International Journal of Computer Processing of Languages
  • Rattasit Sukhahuta +1
  • Conference Article
  • Citations2

Joint Extraction Model of Entity Relations Based on BERT-CRF

  • Oct 01, 2022
  • Yuntao Wang +2
  • Research Article
  • Citations19

A novel prompting method for few-shot NER via LLMs

  • Aug 24, 2024
  • Natural Language Processing Journal
  • Qi Cheng +5
  • Research Article
  • Citations34

A region-based hypergraph network for joint entity-relation extraction

  • Jul 10, 2021
  • Knowledge-Based Systems
  • Qian Wan +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.