• Home
  • Search
  • Cross-lingual dependency parsing for a language with a unique script
  • Cite Icon4
  • https://doi.org/10.1017/nlp.2024.21Copy DOI Icon

Cross-lingual dependency parsing for a language with a unique script

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Abstract Syntactic parsing is one of the areas in Natural Language Processing. The development of large-scale multilingual language models has enabled cross-lingual parsing approaches, which allows us to develop parsers for languages that do not have treebanks available. However, these approaches rely on the assumption that languages share orthographic representations and lexical entries. In this article, we investigate methods for developing a dependency parser for Xibe, a low-resource language that is written in a unique script. We first investigate lexicalized monolingual dependency parsing experiments to examine the effectiveness of word, part-of-speech, and character embeddings as well as pre-trained language models. Results show that character embeddings can significantly improve performance, while pre-trained language models decrease performance since they do not recognize the Xibe script. We also train delexicalized monolingual models, which yield competitive results to the best lexicalized model. Since the monolingual models are trained on a very small training set, we also investigate lexicalized and delexicalized cross-lingual models. We use six closely related languages as source language, which cover a wide range of scripts. In this setting, the delexicalized models achieve higher performance than lexicalized models. A final experiment shows that we can increase performance of the cross-lingual model by combining source languages and selecting the most similar sentences to Xibe as training set. However, all cross-lingual parsing results are still considerably lower than the monolingual model. We attribute the low performance of cross-lingual methods to syntactic and annotation differences as well as to the impoverished input of Universal Dependency Part-of-Speech tags that the delexicalized model has access to.

Similar Papers
  • Conference Article
  • Citations16

Unsupervised Neural Machine Translation for English to Kannada Using Pre-Trained Language Model

  • Oct 03, 2022
  • Shailashree K Sheshadri +4
  • PDF
  • Research Article
  • Citations16

BioBERTurk: Exploring Turkish Biomedical Language Model Development Strategies in Low-Resource Setting.

  • Sep 19, 2023
  • Journal of healthcare informatics research
  • Hazal Türkmen +4
  • PDF
  • Conference Article
  • Citations22

Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

  • Jan 01, 2023
  • Fan Zhou +4
  • Research Article
  • Citations338

Pre-Trained Language Models and Their Applications

  • Jun 01, 2023
  • Engineering
  • Haifeng Wang +4
  • PDF
  • Research Article
  • Citations4

A new computationally efficient method to tune BERT networks – transfer learning

  • Sep 01, 2023
  • Journal of Physics: Conference Series
  • Zian Wang
  • Research Article
  • Citations18

Can Pretrained English Language Models Benefit Non-English NLP Systems in Low-Resource Scenarios?

  • Jan 01, 2024
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Zewen Chi +5
  • PDF
  • Research Article
  • Citations10

AGI-P: A Gender Identification Framework for Authorship Analysis Using Customized Fine-Tuning of Multilingual Language Model

  • Jan 01, 2024
  • IEEE Access
  • Raheem Sarwar +6
  • Research Article
  • Citations27

Natural language processing applications for low-resource languages

  • Feb 28, 2025
  • Natural Language Processing
  • Partha Pakray +2
  • Conference Article
  • Citations3

Chinese-Korean Weibo Sentiment Classification Based on Pre-trained Language Model and Transfer Learning

  • May 06, 2022
  • Hengxuan Wang +3
  • PDF
  • Research Article
  • Citations18

Pre-trained transformer-based language models for Sundanese

  • Apr 13, 2022
  • Journal of Big Data
  • Wilson Wongso +2
  • Research Article
  • Citations13

JointMatcher: Numerically-aware entity matching using pre-trained language models with attention concentration

  • May 16, 2022
  • Knowledge-Based Systems
  • Chen Ye +6
  • Video Transcripts

Can Pre-trained Language Models Interpret Similes as Smart as Human?

  • May 11, 2022
  • Underline Science Inc.
  • Qianyu He +4
  • Research Article
  • Citations77

Bangla-BERT: Transformer-Based Efficient Model for Transfer Learning and Language Understanding

  • Jan 01, 2022
  • IEEE Access
  • M Kowsher +5
  • PDF
  • Conference Article
  • Citations31

PANLP at MEDIQA 2019: Pre-trained Language Models, Transfer Learning and Knowledge Distillation

  • Jan 01, 2019
  • Wei Zhu +6
  • PDF
  • Conference Article
  • Citations8

ANNA”:" Enhanced Language Representation for Question Answering

  • Jan 01, 2022
  • Changwook Jun +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.