• Home
  • Search
  • Triplet Knowledge Distillation Networks for Model Compression
  • https://doi.org/10.1109/ijcnn52387.2021.9534441Copy DOI Icon

Triplet Knowledge Distillation Networks for Model Compression

  • Jul 18, 2021
  • Jialiang Tang +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Knowledge distillation is a widely used neural network model compression technique. In general, the knowledge distillation transfer the knowledge from a large pre-trained teacher network with superior performance to a small student network enables the student network to achieve better performance. This paper proposes a triplet knowledge distillation framework (abbreviated as TKD), which introduces a smaller assistant network into the knowledge distillation structure. The performance of the assistant network is lower than that of the student network. During the training of the TKD, by minimizing the Mean Squared Error(MSE) loss function, the output of the student network will closer to the output of the teacher network and further from that of the assistant network. Therefore, the student network can learn more expressive knowledge from the teacher network while throwing away mistaken knowledge in the assistant network. Finally, the student network achieves a surprising performance even superior to the teacher network. We have demonstrated the effectiveness of TKD by extensive experiments on benchmark datasets(CIFAR-10, CIFAR-100, SVHN, STL-10). When using VGGNet as an experimental model, the student network VGGNet13 achieving 94.29%, 75.30%, 95.53%, and 87.61% accuracy on the CIFAR-10, CIFAR-100, SVHN, and STL-10 datasets, improved by 1.24%, 2.81%, 0.40%, and 2.32%, respectively.

Similar Papers
  • Research Article
  • Citations8

Cosine similarity-guided knowledge distillation for robust object detectors

  • Aug 14, 2024
  • Scientific Reports
  • Sangwoo Park +2
  • Research Article
  • Citations14

Knowledge distillation guided by multiple homogeneous teachers

  • May 31, 2022
  • Information Sciences
  • Quanzheng Xu +2
  • Conference Article
  • Citations1

Medical Image Segmentation Approach via Transformer Knowledge Distillation

  • Apr 28, 2023
  • Tian-Tian Zhang +3
  • Research Article

Distilling Structural Knowledge for Platform-Aware Semantic Segmentation

  • May 01, 2024
  • Journal of Physics: Conference Series
  • Guilin Li +2
  • Research Article
  • Citations31

Skill-Transferring Knowledge Distillation Method

  • Nov 01, 2023
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Shunzhi Yang +5
  • Research Article

Knowledge Distillation for Image Signal Processing Using Only the Generator Portion of a GAN

  • Nov 20, 2022
  • Electronics
  • Youngjun Heo +1
  • PDF
  • Research Article
  • Citations25

Semantic-aware knowledge distillation with parameter-free feature uniformization

  • May 08, 2023
  • Visual Intelligence
  • Guangyu Guo +4
  • Research Article
  • Citations27

A tucker decomposition based knowledge distillation for intelligent edge applications

  • Dec 29, 2020
  • Applied Soft Computing
  • Cheng Dai +3
  • PDF
  • Research Article
  • Citations67

Teacher-student collaborative knowledge distillation for image classification

  • May 04, 2022
  • Applied Intelligence
  • Chuanyun Xu +5
  • Conference Article
  • Citations506

Distilling Knowledge via Knowledge Review

  • Jun 01, 2021
  • Pengguang Chen +3
  • Conference Article
  • Citations5

ShuffleCount: Task-Specific Knowledge Distillation for Crowd Counting

  • Sep 19, 2021
  • Minyang Jiang +2
  • Conference Article
  • Citations249

Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation

  • Aug 01, 2021
  • Taehyeon Kim +4
  • Research Article

Self-Knowledge Distillation and Its Application: A Survey

  • Jan 01, 2026
  • IEEE Transactions on Knowledge and Data Engineering
  • Kai Xu +3
  • PDF
  • Research Article
  • Citations1

Deep Collaborative Learning for Randomly Wired Neural Networks

  • Jul 13, 2021
  • Electronics
  • Ehab Essa +1
  • PDF
  • Research Article
  • Citations3

The Effectiveness of the Squared Error and Higgins-Tsokos Loss Functions on the Bayesian Reliability Analysis of Software Failure Times under the Power Law Process

  • Jan 01, 2019
  • Engineering
  • Freeh N Alenezi +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.