• Home
  • Search
  • Multi-Teacher Distillation With Single Model for Neural Machine Translation
  • Cite Icon21
  • https://doi.org/10.1109/taslp.2022.3153264Copy DOI Icon

Multi-Teacher Distillation With Single Model for Neural Machine Translation

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Knowledge distillation (KD) is an effective strategy for neural machine translation (NMT) to improve the performance of a student model. Usually, the teacher can guide the student to be better by distilling the soft label or data knowledge from the teacher itself. However, the data diversity and teacher knowledge are limited with only one teacher model. Though a natural solution is to adopt multiple randomized teacher models, one big shortcoming is that the model parameters and training costs are largely increased with the number of teacher models. In this work, we explore to mimic multiple teacher distillation from the sub-network space and permuted variants of one single teacher model. Specifically, we train a teacher by multiple sub-network extraction paradigms: sub-layer reordering, layer-drop, and dropout variants. In doing so, one teacher model can provide multiple outputs variants and causes neither additional parameters nor much extra training cost. Experiments on <formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex>$8$</tex></formula> IWSLT datasets: (IWSLT14 En <formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex>$\leftrightarrow$</tex></formula> De, En <formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex>$\leftrightarrow$</tex></formula> Es, and IWSLT17 En <formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex>$\leftrightarrow$</tex></formula> Fr, En <formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex>$\leftrightarrow$</tex></formula> Zh) and the large WMT14 EN <formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex>$\to$</tex></formula> DE translation tasks show that our method even achieves nearly comparable performance with multiple teacher models with different randomized parameters, both word-level, and sequence-level knowledge distillation. Our code is available at GitHub\footnote{https://github.com/dropreg/RLD}

Similar Papers
  • Video Transcripts

Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine Translation

  • May 11, 2022
  • Underline Science Inc.
  • Hongji Wang +5
  • Research Article
  • Citations10

Translating with Bilingual Topic Knowledge for Neural Machine Translation

  • Jul 17, 2019
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Xiangpeng Wei +4
  • PDF
  • Conference Article
  • Citations12

The Unreasonable Volatility of Neural Machine Translation Models

  • Jan 01, 2020
  • Marzieh Fadaee +1
  • Conference Article
  • Citations4

NICT‘s Submission To WAT 2020: How Effective Are Simple Many-To-Many Neural Machine Translation Models?

  • Jan 01, 2020
  • Raj Dabre +1
  • PDF
  • Conference Article
  • Citations854

Sequence-Level Knowledge Distillation

  • Jan 01, 2016
  • Yoon Kim +1
  • Research Article
  • Citations1

Modeling hypotactic structure for Chinese-English neural machine translation of complex sentences

  • Dec 16, 2021
  • Journal of Intelligent &amp; Fuzzy Systems
  • Guoyi Miao +5
  • Supplementary Content
  • Citations8

English-Chinese Machine Translation Based on Transfer Learning and Chinese-English Corpus.

  • Sep 27, 2022
  • Computational Intelligence and Neuroscience
  • Bo Xu
  • Conference Article
  • Citations3

Neural Machine Translation for English-Assamese Language Pair using Transformer

  • Oct 07, 2022
  • Rudra Dutt +3
  • PDF
  • Conference Article
  • Citations365

What do Neural Machine Translation Models Learn about Morphology?

  • Jan 01, 2017
  • Yonatan Belinkov +4
  • Conference Article
  • Citations1

Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation

  • Aug 01, 2024
  • Hyowon Wi +2
  • Research Article
  • Citations1

Weight Saliency search with Semantic Constraint for Neural Machine Translation attacks

  • May 29, 2024
  • Pattern Recognition Letters
  • Wen Han +4
  • Conference Article

A Review of Discourse-level Machine Translation

  • Jan 01, 2020
  • Xiaojun Zhang
  • Research Article

Comparison of Deep Learning Approaches for Conversion of International Classification of Diseases Codes to the Abbreviated Injury Scale

  • Mar 22, 2024
  • medRxiv
  • Ayush Doshi +4
  • Research Article
  • Citations121

Asynchronous Bidirectional Decoding for Neural Machine Translation

  • Apr 27, 2018
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Xiangwen Zhang +5
  • PDF
  • Conference Article
  • Citations12

The LAIX Systems in the BEA-2019 GEC Shared Task

  • Jan 01, 2019
  • Ruobing Li +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.