• Home
  • Search
  • AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM Serving
  • https://doi.org/10.1109/hpca68181.2026.11408539Copy DOI Icon

AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM Serving

  • Jan 31, 2026
  • Xinkai Wang +11 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Generative AI, especially LLM, is driving a fundamental shift in software paradigms, prompting cloud providers to build more efficient serving infrastructures. To meet the computational demands of emerging software, modern CPU processors are integrating Accelerator Units (AU) in the pipeline to accelerate key operations, such as Intel AMX for matrix multiplication. Current practices that dedicate AU-enabled CPU exclusively to LLM serving lead to significant resource waste and inferior efficiency. To this end, sharing AU-enabled CPU with general workloads is necessary to harvest redundant resources and improve platform performance-per-watt. However, perfectly sharing AU can be challenging since they introduce three-dimensional variations: variable usage patterns, compulsory frequency interferences, and dissimilar resource bounds. Existing resource managers are oblivious to complex Accelerator Unit Variations (AUV), resulting in performance and efficiency degradations of up to 50 % in shared environments. Therefore, this paper introduces AUM, a novel AU-aware resource manager designed to handle AUV and maximize the efficiency of shared processors. AUM has two cooperative components with three stages for three-dimensional AUV. The background profiler characterizes the usage, frequency, and resource information into a discrete model, guiding the runtime controller to analyze usage-aware requirements, select frequency-aware divisions, and make bound-aware resource decisions. Through extensive evaluations on production AU-enabled CPUs, we show that AUM improves CPU efficiency by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$4.7-8.8 \%$</tex> while maintaining high-performance AU applications by reducing SLO violations by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{7 - 1 1 \%}$</tex> compared with state-of-the-art resource managers.

Similar Papers
  • Conference Article
  • Citations3

Resource management in the Cronus distributed operating system

  • Aug 01, 1987
  • R Schantz +2
  • Conference Article
  • Citations1

Adaptive fair resource management with an arbiter for multi-tier computing systems

  • Sep 01, 2009
  • Naoki Hayashi +2
  • Research Article
  • Citations1

New Directions in Resource Recreation Management: A Response to the Educational Challenge

  • Apr 01, 1997
  • SCHOLE: A Journal of Leisure Studies and Recreation Education
  • Dave Robinson +3
  • Book Chapter
  • Citations1

Chapter 8 - Two-Phase Commit

  • Jan 01, 2009
  • Principles of Transaction Processing
  • Philip A Bernstein +1
  • Conference Article
  • Citations12

Indeterminacy, monitors, and dataflow

  • Jan 01, 1977
  • Arvind +2
  • Research Article
  • Citations1

Ethical and Social Risk Awareness in Generative AI (GenAI): The Role of Mindset and GenAI Literacy

  • Aug 30, 2025
  • International Journal of Technology in Education
  • Sahin Gokcearslan +4
  • Preprint Article

Leveraging Generative AI for Real-Time Anomaly Detection in SAP FICO: A Paradigm Shift in Financial Governance

  • Jul 31, 2025
  • Srinivas Raju Gottimukkala
  • Research Article
  • Citations16

The Identification of Network Intrusions with Generative Artificial Intelligence Approach for Cybersecurity

  • Oct 09, 2024
  • Journal of Web Applications and Cyber Security
  • Himanshu Sinha
  • Conference Article
  • Citations259

PARTIES

  • Apr 04, 2019
  • Shuang Chen +2
  • Conference Article
  • Citations2

Modular reasoning about open systems: a case study of distributed commit

  • Dec 06, 1993
  • Ranjini Das +1
  • Single Report

Multi-static Serial LiDAR for Surveillance and Identification of Marine Life at MHK Installations

  • Jun 30, 2017
  • Gabriel Alsenas +2
  • Research Article
  • Citations1

Multicausality as an Increase in the Latency Period for the Diagnosis of Parkinson's Disease in Primary Care: A Case Report

  • May 01, 2021
  • Journal of Clinical and Medical Research
  • Hector Riquelme-Heras
  • Book Chapter

Distribution and Logistics Network Design in Industry 5.0: Models, Methods, and the Impact of Generative AI and Large Language Models

  • Oct 15, 2025
  • Yu Hao +3
  • Research Article
  • Citations87

Generative AI: is it a paradigm shift for higher education?

  • Mar 22, 2024
  • Studies in Higher Education
  • Xianghan O’Dea
  • Single Report

Augmented Human Intelligence: Converging Generative AI, Quantum Computing, and XR for Enhanced Human-Machine Synergy

  • Mar 21, 2025
  • Murali Krishna Pasupuleti
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.