• Home
  • Search
  • Clustering and Bootstrapping Based Framework for News Knowledge Base Completion
  • Cite Icon1
  • https://doi.org/10.31577/cai_2021_2_318Copy DOI Icon

Clustering and Bootstrapping Based Framework for News Knowledge Base Completion

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Extracting the facts, namely entities and relations, from unstructured sources is an essential step in any knowledge base construction. At the same time, it is also necessary to ensure the completeness of the knowledge base by incrementally extracting the new facts from various sources. To date, the knowledge base completion is studied as a problem of knowledge refinement where the missing facts are inferred by reasoning about the information already present in the knowledge base. However, facts missed while extracting the information from multilingual sources are ignored. Hence, this work proposed a generic framework for knowledge base completion to enrich a knowledge base of crime-related facts extracted from online news articles in the English language, with the facts extracted from low resourced Indian language Hindi news articles. Using the framework, information from any low-resourced language news articles can be extracted without using language-specific tools like POS tags and using an appropriate machine translation tool. To achieve this, a clustering algorithm is proposed, which explores the redundancy among the bilingual collection of news articles by representing the clusters with knowledge base facts unlike the existing Bag of Words representation. From each cluster, the facts extracted from English language articles are bootstrapped to extract the facts from comparable Hindi language articles. This way of bootstrapping within the cluster helps to identify the sentences from a low-resourced language that are enriched with new information related to the facts extracted from a high-resourced language like English. The empirical result shows that the proposed clustering algorithm produced more accurate and high-quality clusters for monolingual and cross-lingual facts, respectively. Experiments also proved that the proposed framework achieves a high recall rate in extracting the new facts from Hindi news articles.

Similar Papers
  • Book Chapter
  • Citations11

Knowledge Base Completion Using Matrix Factorization

  • Jan 01, 2015
  • Wenqiang He +3
  • Research Article

Part-of-Speech (POS) Tagging of Low-Resource Language (Limbu) with Deep learning

  • Nov 13, 2024
  • Panamerican Mathematical Journal
  • Abigail Rai
  • Book Chapter
  • Citations10

Evaluating Language Models for Knowledge Base Completion

  • Jan 01, 2023
  • Blerta Veseli +3
  • PDF
  • Research Article
  • Citations4

Enabling personalised disease diagnosis by combining a patient’s time-specific gene expression profile with a biomedical knowledge base

  • Feb 07, 2024
  • BMC Bioinformatics
  • Ghanshyam Verma +2
  • Research Article
  • Citations33

Fake news detection in low-resource languages: A novel hybrid summarization approach

  • May 02, 2024
  • Knowledge-Based Systems
  • Jawaher Alghamdi +2
  • Conference Article

Fact Discovery from Knowledge Base via Facet Decomposition

  • Jan 01, 2019
  • Zihao Fu +3
  • Research Article
  • Citations7

Co-occurrence Weight Selection in Generation of Word Embeddings for Low Resource Languages

  • Jan 09, 2019
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Veysel Yücesoy +1
  • Research Article
  • Citations7

Stop the Hate, Spread the Hope: An Ensemble Model for Hope Speech Detection in English and Dravidian Languages

  • Feb 12, 2025
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Deepawali Sharma +3
  • Research Article
  • Citations10

Heuristics for Reconciling Independent Knowledge Bases

  • Sep 01, 1993
  • Information Systems Research
  • Andrew Trice +1
  • Supplementary Content
  • Citations22

Automatic Extraction of Facts, Relations, and Entities for Web-Scale Knowledge Base Population

  • Jan 01, 2012
  • Publications of the UdS (Saarland University)
  • Ndapandula Nakashole
  • Conference Article
  • Citations13

Automatic Entity Recognition and Typing in Massive Text Data

  • Jun 26, 2016
  • Xiang Ren +3
  • Conference Article
  • Citations3

Knowledge Base Completion with transfer learning using BERT and fastText

  • Oct 19, 2022
  • Thuy-Anh Nguyen Thi +5
  • Research Article
  • Citations33

Knowledge Base Completion by Variational Bayesian Neural Tensor Decomposition

  • Jun 26, 2018
  • Cognitive Computation
  • Lirong He +5
  • Conference Article

Sentiment Analysis for Low Resource Language Using Machine Learning and Deep Learning

  • Nov 21, 2025
  • AIJR Proceedings
  • Tajinder Kaur +1
  • Research Article
  • Citations7

Impact of Similarity Measures in Graph-based Automatic Text Summarization of Konkani Texts

  • Feb 21, 2023
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Jovi D'Silva +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.