• Home
  • Search
  • Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey.
  • https://doi.org/10.1109/tpami.2026.3676982Copy DOI Icon

Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract

The rapid advancement of large language models (LLMs) has intensified the need for effective mechanisms to transform continuous multimodal data into discrete representations suitable for language-based processing. Discrete tokenization, with vector quantization (VQ) as a central approach, offers both computational efficiency and compatibility with LLM architectures. Despite its growing importance, there is a lack of a comprehensive survey that systematically examines VQ techniques in the context of LLM-based systems. This work fills this gap by presenting the first structured taxonomy and analysis of discrete tokenization methods designed for LLMs. We categorize 8 representative VQ variants that span classical and modern paradigms and analyze their algorithmic principles, training dynamics, and integration challenges with LLM pipelines. Beyond algorithm-level investigation, we discuss existing research in terms of classical applications without LLMs, LLM-based single-modality systems, and LLM-based multimodal systems, highlighting how quantization strategies influence alignment, reasoning, and generation performance. In addition, we identify key challenges including codebook collapse, unstable gradient estimation, and modality-specific encoding constraints. Finally, we discuss emerging research directions such as dynamic and task-adaptive quantization, unified tokenization frameworks, and biologically inspired codebook learning. This survey bridges the gap between traditional vector quantization and modern LLM applications, serving as a foundational reference for the development of efficient and generalizable multimodal systems. A continuously updated version is available at: https://github.com/jindongli-Ai/LLM-Discrete-Tokenization-Survey.

Similar Papers
  • Research Article

Research and selection of Large Learning Models for automation of ABAP-code migration

  • Sep 24, 2025
  • Management of Development of Complex Systems
  • Oleg Pozdnyakov +1
  • Research Article
  • Citations10

Does GPT-4 surpass human performance in linguistic pragmatics?

  • Jun 10, 2025
  • Humanities and Social Sciences Communications
  • Ljubiša Bojić +2
  • Research Article
  • Citations1

Optimization of traditional methods for determining the similarity of project names and purchases using large language models

  • Apr 01, 2024
  • Litera
  • Aleksei Aleksandrovich Golikov +2
  • Research Article

A Multimodal AI System: Comparing LLMs and Theorem Proving Systems

  • Feb 21, 2026
  • Electronics
  • Phillip G Bradford +1
  • Research Article
  • Citations51

Large language models for biomedicine: foundations, opportunities, challenges, and best practices.

  • Apr 24, 2024
  • Journal of the American Medical Informatics Association : JAMIA
  • Satya S Sahoo +8
  • Conference Article
  • Citations5

Game Theory Meets Large Language Models: A Systematic Survey

  • Sep 01, 2025
  • Haoran Sun +3
  • Research Article
  • Citations5

Facilitating Large Language Model Russian Adaptation with Learned Embedding Propagation

  • Dec 30, 2024
  • Journal of Language and Education
  • Mikhail Tikhomirov +1
  • Conference Article

Bio-Inspired LLMs Forgetting: Integrating Neuroscience and Computational Mechanisms

  • Oct 17, 2025
  • Yuchen Xie
  • Research Article

FlashDecoding++Next: High Throughput LLM Inference with Latency and Memory Optimization

  • Jan 01, 2025
  • IEEE Transactions on Computers
  • Guohao Dai +10
  • Research Article
  • Citations1

Predicting pediatric diagnostic imaging patient no-show and extended wait-times using LLMs, regression, and tree based models

  • Sep 03, 2025
  • Frontiers in Artificial Intelligence
  • Daniel Rafique +11
  • Research Article
  • Citations10

International Symposium on Ruminant Physiology: Leveraging computer vision, large language models, and multimodal machine learning for optimal decision making in dairy farming.

  • Jul 01, 2025
  • Journal of dairy science
  • Rafael E P Ferreira +1
  • Conference Article

LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system

  • Aug 08, 2025
  • Huanyu Li +3
  • Research Article

Modern Approaches to Using Knowledge Bases to Address the Challenges of Large Language Models

  • May 12, 2025
  • NaUKMA Research Papers. Computer Science
  • Maksym Androshchuk
  • Research Article

Large language models in healthcare and biomedical informatics: A comprehensive review

  • Jan 01, 2026
  • Innovation and Emerging Technologies
  • Andrew Hornback +8
  • PDF
  • Research Article
  • Citations14

LLM-Based Agents for Tool Learning: A Survey

  • Jun 26, 2025
  • Data Science and Engineering
  • Weikai Xu +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.