• Home
  • Search
  • EPMORE: Explainable Process Mixture-of-Experts
  • https://doi.org/10.21203/rs.3.rs-8374807/v2Copy DOI Icon

EPMORE: Explainable Process Mixture-of-Experts

  • Feb 5, 2026
  • Wei Sheng +6 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Abstract Large language models (LLMs) are primarily built on the Transformer architecture, in which all hidden layers share a fixed-dimensional representation spaces. This homogeneity constrains representational capacity, impedes interpretability, and induces computational redundancy. We propose EPMORE (Explainable Process Mixture-of-Experts), a novel architecture that models inference as a process of dimensional elevation, Ensure the entire inference/training process, intermediate states observable and explainable, and ensure the whole process is traceable end-to-end. EPMORE decomposes the entire inference/training process into a hierarchical sequence of representation spaces — from a semantic space (128 dimensions), to one or more logical spaces (512 dimensions each), and finally to fact-expert representation spaces (1024 dimensions) — allowing deeper network stages to encode progressively richer and more abstract features. A core component, Middle Output Reuse (MOR), enables each layer to produce interpretable intermediate predictions. Theoretically, forward propagation can be interpreted as representation-space expansion, while backward propagation corresponds to a dimensional contraction process. Experiments show that, compared with dense and conventional mixture-of-experts (MoE, Deepseek) baselines, EPMORE improves interpretability, activation sparsity, parameter independence, and inference performance while reducing computational cost. These findings suggest that hierarchical dimensional elevation is a promising alternative to standard Transformer design.

Similar Papers
  • Research Article
  • Citations11

Denoising Alignment with Large Language Model for Recommendation

  • Jan 24, 2025
  • ACM Transactions on Information Systems
  • Yingtao Peng +7
  • Conference Article
  • Citations8

Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack

  • Dec 02, 2024
  • CCS 2024 - Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security
  • Guanzhong Chen +6
  • Conference Article

AlphaFuse: Learn ID Embeddings for Sequential Recommendation in Null Space of Language Embeddings

  • Jul 13, 2025
  • Guoqing Hu +5
  • Conference Article

A Pre-training Method Inspired by Large Language Model for Power Named Entity Recognition

  • Nov 24, 2023
  • Qionglan Na +5
  • Preprint Article

From Tokens To Agents: A Researcher's Guide To Understanding Large Language Models

  • Feb 27, 2026
  • arXiv (Cornell University)
  • Daniele Barolo
  • Research Article
  • Citations2

Improving Narrative Coherence in Dense Video Captioning through Transformer and Large Language Models

  • Jun 03, 2025
  • Journal of Innovative Image Processing
  • Dvijesh Bhatt +1
  • Research Article

A systematic review on the generative AI applications in human medical genetics

  • Jan 20, 2026
  • Frontiers in Genetics
  • Anton Changalidis +3
  • Book Chapter

Explainable AI (XAI) for Classical ML and LLMs

  • May 16, 2026
  • Brijendra Parasnath Gupta +6
  • Conference Article
  • Citations1

Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization

  • Sep 01, 2025
  • Xinhao Yao +7
  • Research Article
  • Citations1

Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Kyle Cox +8
  • Preprint Article

Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation

  • Dec 05, 2025
  • arXiv (Cornell University)
  • Pengzhen Xie +1
  • Preprint Article

DeepSeMS: a large language model reveals hidden biosynthetic potential of the global ocean microbiome

  • Apr 16, 2025
  • Research Square
  • Na Jiao +9
  • Research Article
  • Citations2

Dynamic-Width Speculative Beam Decoding for LLM Inference

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Zongyue Qin +4
  • Conference Article

Generative AI for Code Translation: A Systematic Mapping Study

  • Oct 15, 2025
  • Aymane Rgaguena +2
  • Video Transcripts

What Language Model to Train if You Have One Million GPU Hours?

  • May 17, 2022
  • Underline Science Inc.
  • Iz Beltagy +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.