• Home
  • Search
  • FBQuant: FeedBack Quantization for Large Language Models
  • https://doi.org/10.24963/ijcai.2025/844Copy DOI Icon

FBQuant: FeedBack Quantization for Large Language Models

  • Sep 1, 2025
  • Yijiang Liu +6 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Deploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user privacy. However, on-device deployment is challenging due to the limited computational resources of edge devices. In particular, the key bottleneck stems from memory bandwidth constraints related to weight loading. Weight-only quantization effectively reduces memory access, yet often induces significant accuracy degradation. Recent efforts to incorporate sub-branches have shown promise for mitigating quantization errors, but these methods either lack robust optimization strategies or rely on suboptimal objectives. To address these gaps, we propose FeedBack Quantization (FBQuant), a novel approach inspired by negative feedback mechanisms in automatic control. FBQuant inherently ensures that the reconstructed weights remain bounded by the quantization process, thereby reducing the risk of overfitting. To further offset the additional latency introduced by sub-branches, we develop an efficient CUDA kernel that decreases 60% of extra inference time. Comprehensive experiments demonstrate the efficiency and effectiveness of FBQuant across various LLMs. Notably, for 3-bit Llama2-7B, FBQuant improves zero-shot accuracy by 1.2%.

Similar Papers
  • Research Article

English

  • Feb 28, 2025
  • International Journal of Computer Trends and Technology
  • Dhivya Nagasubramanian
  • PDF
  • Research Article
  • Citations3

Exploring the potential of lightweight large language models for AI-based mental health counselling task: a novel comparative study

  • Jul 02, 2025
  • Scientific Reports
  • Ritesh Maurya +4
  • Conference Article
  • Citations9

Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures

  • Oct 27, 2024
  • Ruiyang Qin +11
  • Research Article
  • Citations12

ROFED-LLM: Robust Federated Learning for Large Language Models in Adversarial Wireless Environments

  • Jan 01, 2026
  • IEEE Transactions on Network Science and Engineering
  • Haoyu Wang +6
  • Research Article

Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines.

  • Apr 29, 2026
  • IEEE journal of biomedical and health informatics
  • Guifeng Deng +11
  • Conference Article
  • Citations1

A First Look at LLM-powered Smartphones

  • Oct 27, 2024
  • Liangxuan Wu +4
  • Research Article
  • Citations5

Reflective Dialogues with a Humanoid Robot Integrated with an LLM and a Curated NLU System for Positive Behavioral Change in Older Adults

  • Nov 07, 2024
  • Electronics
  • Ryan Browne +17
  • Conference Article
  • Citations13

Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and Synthesis

  • Jun 23, 2024
  • Ruiyang Qin +7
  • Research Article

LLMGuard : Safeguarding Real-Time Inference for Large Language Models on Edge Devices

  • Feb 25, 2026
  • ACM Transactions on Software Engineering and Methodology
  • Yu Sun +4
  • Research Article
  • Citations5

GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices

  • May 28, 2025
  • Proceedings of the AAAI Symposium Series
  • Mozhgan Navardi +6
  • Conference Article

Generative AI in Assessment and Feedback Generation in Higher Education: A Systematic Review

  • Sep 18, 2025
  • Amir Yavariabdi +3
  • Discussion
  • Citations6

Comparative analysis of large language models on rare disease identification

  • Apr 01, 2025
  • Orphanet Journal of Rare Diseases
  • Guangyu Ao +5
  • Research Article
  • Citations4

VaVLM: Toward Efficient Edge-Cloud Video Analytics With Vision-Language Models

  • Jun 01, 2025
  • IEEE Transactions on Broadcasting
  • Yang Zhang +6
  • Research Article
  • Citations6

Comparative Study on Energy Consumption of Neural Networks by Scaling of Weight-Memory Energy Versus Computing Energy for Implementing Low-Power Edge Intelligence

  • Jul 05, 2025
  • Electronics
  • Ilpyung Yoon +2
  • PDF
  • Research Article
  • Citations60

Ethical Considerations and Fundamental Principles of Large Language Models in Medical Education: Viewpoint.

  • Aug 01, 2024
  • Journal of medical Internet research
  • Li Zhui +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.