• Home
  • Search
  • Building a knowledge graph by using cross-lingual transfer method and distributed MinIE algorithm on apache spark
  • Cite Icon20
  • https://doi.org/10.1007/s00521-020-05495-1Copy DOI Icon

Building a knowledge graph by using cross-lingual transfer method and distributed MinIE algorithm on apache spark

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The simplest and effective way to store human knowledge through centuries was using text. Along with the advancement of technology nowadays, the volume of text has grown to be larger and larger. To extract useful information from this amount of text becomes an exceptionally complex task. As an effort to solve that problem, in this paper, we present a pipeline to extract core knowledge from large quantity text using distributed computing. The components of our pipeline are systems that were known to yield good results. The outputs of our proposed system are stored in a knowledge graph. A knowledge graph is a graph for storing knowledge in the form of triples (head, relation, tail). Some of the existing knowledge graphs in the world are Google knowledge graph, YAGO, DBLP, or DBpedia. These knowledge graphs have one thing in common—they are in English. The English language is studied by many researchers in the world and it had become a rich-resource language (with many natural language processing tools and data set). Vietnamese, on the other hand, is a low-resource language. Therefore, we use cross-lingual transfer method to build a Vietnamese knowledge graph. Firstly, we collect data in form of text about Vietnam tourism, which was written mostly in Vietnamese, using Google search and Wikipedia. In the next step, we translate them into English with Google Translate and use English Natural Language Processing tools like Stanford Parser, Co-referencing, ClausIE, MinIE to extract useful triples from this text. Lastly, the triples are translated back to Vietnamese to build a Vietnam tourism knowledge graph. Since we are working with massive text, we develop a distributed algorithm to extract triples from sentences of massive text. This is a distributed version of MinIE, which was originally developed for a single machine model. In Apache Spark framework, we divide massive text into many smaller parts and move them to the worker nodes with distributed MinIE function. Spark distributed MinIE will extract the triples of sentences in the local text of this worker node in parallel. Finally, the result of worker nodes will be sent back to the master node for building the knowledge graph. We conduct experiments with the distributed MinIE on spark cluster to prove the outperformance of our proposed algorithm.

Similar Papers
  • Research Article
  • Citations8

Learning knowledge graph embedding with a bi-directional relation encoding network and a convolutional autoencoder decoding network

  • Jan 07, 2021
  • Neural Computing and Applications
  • Kairong Hu +4
  • Book Chapter
  • Citations2

Creation of Knowledge Graph for Client Complaint Management System

  • Jan 01, 2021
  • Shreya Shinde +4
  • PDF
  • Research Article
  • Citations2

A unified approach to publish semantic annotations of agricultural documents as knowledge graphs

  • Jul 06, 2024
  • Smart Agricultural Technology
  • Nadia Yacoubi Ayadi +9
  • Research Article
  • Citations6

An Ontology-Based Knowledge Methodology in the Medical Domain in the Latin America: the Study Case of Republic of Panama.

  • Jan 01, 2018
  • Acta Informatica Medica
  • Denis Cedenomoreno +1
  • Research Article
  • Citations26

Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs

  • Oct 08, 2024
  • Proceedings of the ACM on Programming Languages
  • Federico Cassano +9
  • Research Article
  • Citations16

Integration of Spark framework in Supply Chain Management

  • Jan 01, 2016
  • Procedia Computer Science
  • Harjeet Singh Jaggi +1
  • Research Article

Abstract 6391: Scalable colorectal cancer (CRC) diagnosis model with arterial and unenhanced phases of KRAS mutational status using Apache Spark

  • Jun 15, 2022
  • Cancer Research
  • Mary Adetutu Adewunmi
  • Research Article

Belonging and Mobility: Constructing Sentiment-embedded Event Evolutionary Graph for Diaspora Oral Archives

  • Jan 22, 2026
  • Journal on Computing and Cultural Heritage
  • Jing Zhou +1
  • Single Report
  • Citations1

Making Semantic Structures Explicit: Developing and Evaluating Tools and Techniques to Support Understanding of Large Cybersecurity Corpora

  • Jan 01, 2022
  • Ira Monarch +5
  • Research Article
  • Citations4

The Knowledge Graph for Macroeconomic Analysis with Alternative Big Data

  • Jan 01, 2020
  • SSRN Electronic Journal
  • Yucheng Yang +3
  • Research Article
  • Citations38

DI-Mondrian: Distributed improved Mondrian for satisfaction of the L-diversity privacy model using Apache Spark

  • Aug 09, 2020
  • Information Sciences
  • Farough Ashkouti +2
  • Conference Article

Multimodal Methods for Improving Natural Language Processing in Low-Resource Languages:Survey

  • Oct 30, 2025
  • Ranganath Kanakam +3
  • Research Article
  • Citations7

Sentiment analysis and prediction of polarity vaccines based on Twitter data using deep NLP techniques

  • Nov 29, 2022
  • Radioelectronic and Computer Systems
  • Hassan Badi +4
  • Conference Article
  • Citations35

SparkSW: Scalable Distributed Computing System for Large-Scale Biological Sequence Alignment

  • May 01, 2015
  • Guoguang Zhao +2
  • Research Article

Event-Argument Linking in Disaster Domain

  • Jan 01, 2022
  • IEEE Access
  • Sovan Kumar Sahoo +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.