• Home
  • Search
  • A Practical Tensor-Network Compression Pipeline for Production-Scale Large Language Models
  • https://doi.org/10.48550/arxiv.2602.01613Copy DOI Icon

A Practical Tensor-Network Compression Pipeline for Production-Scale Large Language Models

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Large language models are limited in deployment by GPU memory and inference latency. We present Minima, a production compression pipeline that learns where and how to structurally compress a Transformer and turns that compression into real serving gains. Minima trains a lightweight convolutional predictor to estimate layer- and patch-level sensitivity, applies a mixture of Tucker, tensor-train, and tensor-ring decompositions to low-sensitivity regions, performs a short healing fine-tune, and executes the resulting operators with custom Triton and CUDA kernels. The reduced memory footprint enables speculative decoding with a small draft model and a larger verifier. On Qwen3-32B at an 8k-token context window, Minima reduces peak VRAM from 64 GiB to 40 GiB. For a single active request, throughput increases from 40 tokens per second (baseline) to 50 tokens per second (Minima) and 75 tokens per second (Minima with speculative decoding). Under 50 parallel requests, throughput is 34, 44, and 53 tokens per second respectively, showing that Minima remains effective under high concurrency even when speculative decoding gains compress. We position Minima relative to recent tensor-network, low-rank plus quantization, and cross-layer sharing methods, and argue that it is a practical step toward more aggressive structural compression via shared tensor backbones with tiny per-layer adapters.

Similar Papers
  • Research Article

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

  • Mar 14, 2026
  • Qiyang Li +7
  • Conference Article
  • Citations1

ELEC: Efficient Large Language Model-Empowered Click-Through Rate Prediction

  • Jul 13, 2025
  • Rui Dong +2
  • Research Article

Simulation Study on Real-Time Autonomous Driving Decision-Making Using BEV Perception and Large Language Models

  • Mar 10, 2026
  • Technologies
  • Gaosong Shi +2
  • Research Article

Optimizing Distributed LLM Inference for Heterogeneous Workers through Dynamic Graph Partitioning

  • Mar 01, 2026
  • reposiTUm (TU Wien)
  • Gabriel Kitzberger
  • Research Article

Aggressive Speculative Decoding for Efficient Chain-of-Thought Reasoning

  • Dec 01, 2025
  • Tsinghua Science & Technology
  • Huanran Zheng +3
  • Research Article

Latency Adjustable Transformer Encoder for Language Understanding.

  • Jun 01, 2025
  • IEEE transactions on neural networks and learning systems
  • Sajjad Kachuee +1
  • Conference Article

RankMixer: Scaling Up Ranking Models in Industrial Recommenders

  • Nov 08, 2025
  • Jie Zhu +20
  • Conference Article
  • Citations13

PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

  • Jan 01, 2024
  • Dongjie Yang +5
  • Research Article

LLMOps for Streaming Data: Bridging NLP and Event Pipelines

  • Sep 30, 2025
  • International Journal of Science and Research Archive
  • Devarsh Hemantbhai Patel
  • Conference Article
  • Citations1

Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG Systems

  • Jul 07, 2025
  • Hyungwoo Lee +7
  • Preprint Article

Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems

  • Feb 13, 2026
  • arXiv (Cornell University)
  • Sadia Asif +1
  • Research Article
  • Citations10

PQCache: Product Quantization-based KVCache for Long Context LLM Inference

  • Jun 17, 2025
  • Proceedings of the ACM on Management of Data
  • Hailin Zhang +7
  • Research Article

Enabling efficient low-bit quantization based on matrix product operators for KV cache compression.

  • May 01, 2026
  • Neural networks : the official journal of the International Neural Network Society
  • Jia-Qi Wang +5
  • Research Article

LLMGuard : Safeguarding Real-Time Inference for Large Language Models on Edge Devices

  • Feb 25, 2026
  • ACM Transactions on Software Engineering and Methodology
  • Yu Sun +4
  • Conference Article

FBQuant: FeedBack Quantization for Large Language Models

  • Sep 01, 2025
  • Yijiang Liu +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.