• Home
  • Search
  • SensiMix: Sensitivity-Aware 8-bit index & 1-bit value mixed precision quantization for BERT compression
  • Open Access IconOpen Access
  • Cite Icon1
  • https://doi.org/10.1371/journal.pone.0265621.r006Copy DOI Icon

SensiMix: Sensitivity-Aware 8-bit index & 1-bit value mixed precision quantization for BERT compression

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Given a pre-trained BERT, how can we compress it to a fast and lightweight one while maintaining its accuracy? Pre-training language model, such as BERT, is effective for improving the performance of natural language processing (NLP) tasks. However, heavy models like BERT have problems of large memory cost and long inference time. In this paper, we propose SensiMix (Sensitivity-Aware Mixed Precision Quantization), a novel quantization-based BERT compression method that considers the sensitivity of different modules of BERT. SensiMix effectively applies 8-bit index quantization and 1-bit value quantization to the sensitive and insensitive parts of BERT, maximizing the compression rate while minimizing the accuracy drop. We also propose three novel 1-bit training methods to minimize the accuracy drop: Absolute Binary Weight Regularization, Prioritized Training, and Inverse Layer-wise Fine-tuning. Moreover, for fast inference, we apply FP16 general matrix multiplication (GEMM) and XNOR-Count GEMM for 8-bit and 1-bit quantization parts of the model, respectively. Experiments on four GLUE downstream tasks show that SensiMix compresses the original BERT model to an equally effective but lightweight one, reducing the model size by a factor of 8× and shrinking the inference time by around 80% without noticeable accuracy drop.

Similar Papers
  • Video Transcripts

Can Pre-trained Language Models Interpret Similes as Smart as Human?

  • May 11, 2022
  • Underline Science Inc.
  • Qianyu He +4
  • PDF
  • Research Article
  • Citations7

APRE: Annotation-Aware Prompt-Tuning for Relation Extraction

  • Feb 21, 2024
  • Neural Processing Letters
  • Chao Wei +5
  • Video Transcripts

Lexicon-Based Graph Convolutional Network for Chinese Word Segmentation

  • Oct 23, 2021
  • Underline Science Inc.
  • Kaiyu
  • Research Article
  • Citations20

A Survey on Automatic Generation of Figurative Language: From Rule-based Systems to Large Language Models

  • May 14, 2024
  • ACM Computing Surveys
  • Huiyuan Lai +1
  • Research Article

TOWARDS CROSS-ATTENTION PRE-TRAINING IN NEURAL MACHINE TRANSLATION

  • Oct 31, 2022
  • Tạp chí Khoa học
  • Khang Pham
  • PDF
  • Research Article
  • Citations16

BioBERTurk: Exploring Turkish Biomedical Language Model Development Strategies in Low-Resource Setting.

  • Sep 19, 2023
  • Journal of healthcare informatics research
  • Hazal Türkmen +4
  • PDF
  • Research Article
  • Citations35

Unrestricted Attention May Not Be All You Need–Masked Attention Mechanism Focuses Better on Relevant Parts in Aspect-Based Sentiment Analysis

  • Jan 01, 2022
  • IEEE Access
  • Ao Feng +2
  • PDF
  • Conference Article
  • Citations22

Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

  • Jan 01, 2023
  • Fan Zhou +4
  • Research Article

Task-Adaptive and Multi-Level Contextual Understanding for Emotion Recognition in Conversations

  • Feb 09, 2026
  • Applied Sciences
  • Xiaomeng Yao +4
  • Research Article
  • Citations3

Exploring Named Entity Recognition via MacBERT-BiGRU and Global Pointer with Self-Attention

  • Dec 03, 2024
  • Big Data and Cognitive Computing
  • Chengzhe Yuan +6
  • Research Article
  • Citations20

Prompt for extraction: Multiple templates choice model for event extraction

  • Feb 20, 2024
  • Knowledge-Based Systems
  • Jiaren Peng +3
  • PDF
  • Research Article
  • Citations89

From Word Embeddings to Pre-Trained Language Models: A State-of-the-Art Walkthrough

  • Sep 01, 2022
  • Applied Sciences
  • Mourad Mars
  • Book Chapter
  • Citations5

Accelerating Pretrained Language Model Inference Using Weighted Ensemble Self-distillation

  • Jan 01, 2021
  • Jun Kong +2
  • Research Article
  • Citations2

A Unified Framework of Medical Information Annotation and Extraction for Chinese Clinical Text

  • Jan 01, 2022
  • SSRN Electronic Journal
  • Enwei Zhu +3
  • Research Article
  • Citations1

Mixed Information Bottleneck for Location Metonymy Resolution Using Pre-trained Language Models

  • Nov 08, 2025
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Hao Wang +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.