• Home
  • Search
  • Arlo: Serving Transformer-based Language Models with Dynamic Input Lengths
  • Cite Icon1
  • https://doi.org/10.1145/3673038.3673124Copy DOI Icon

Arlo: Serving Transformer-based Language Models with Dynamic Input Lengths

  • Aug 12, 2024
  • Xin Tan +4 more
Show More
  • Abstract
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

A prominent challenge in serving requests for NLP tasks is handling the varying length of input texts. Existing solutions, such as uniform zero-padding and compiler support, suffer from either computational inefficiency or suboptimal latency. To address these practical issues, we propose an approach called polymorphing. Polymorphing involves creating and utilizing multiple runtimes of the model, each statically compiled with a different input length, to serve requests accordingly. This fine-grained use of statically-compiled runtimes reduces the overheads of zero-padding while improving latency performance compared to dynamic compilation. To practically realize polymorphing, we have developed an inference scheduling system, Arlo, which leverages the observed input length distribution to periodically allocate compute resources across multiple runtimes by solving an integer linear program. Upon request arrival, Arlo uses a multi-level queue-based heuristic to dispatch requests to the most suitable runtime instances, efficiently adapting to the dynamics of request length and instance load. Extensive testbed evaluations and large-scale simulations using production traces demonstrate Arlo's promising potential. It achieves 23.7%–98.1% mean latency reductions compared to existing schemes while significantly reducing tail latency.

Loading PDF

Similar Papers
  • Research Article
  • Citations7

The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Ke Shen
  • PDF
  • Research Article
  • Citations26

Predicting Generalized Anxiety Disorder From Impromptu Speech Transcripts Using Context-Aware Transformer-Based Neural Networks: Model Evaluation Study

  • Mar 28, 2023
  • JMIR Mental Health
  • Bazen Gashaw Teferra +1
  • Research Article
  • Citations14

Fine-Tuned Understanding: Enhancing Social Bot Detection With Transformer-Based Classification

  • Jan 01, 2024
  • IEEE Access
  • Amine Sallah +6
  • PDF
  • Research Article
  • Citations47

MuLan-Methyl—multiple transformer-based language models for accurate DNA methylation prediction

  • Dec 28, 2022
  • GigaScience
  • Wenhuan Zeng +2
  • Research Article
  • Citations28

Reward modeling for mitigating toxicity in transformer-based language models

  • Jul 20, 2022
  • Applied Intelligence
  • Farshid Faal +2
  • Research Article
  • Citations77

Bangla-BERT: Transformer-Based Efficient Model for Transfer Learning and Language Understanding

  • Jan 01, 2022
  • IEEE Access
  • M Kowsher +5
  • PDF
  • Conference Article
  • Citations16

Time-Stamped Language Model: Teaching Language Models to Understand The Flow of Events

  • Jan 01, 2021
  • Hossein Rajaby Faghihi +1
  • Conference Article

Modeling cognitive processes of natural reading with transformer-based Language Models

  • Dec 10, 2024
  • Bruno Bianchi +2
  • Research Article

Axiom Generation for Automated Ontology Construction from Texts Through Schema Mapping

  • Jan 26, 2026
  • Machine Learning and Knowledge Extraction
  • Tsitsi Zengeya +2
  • PDF
  • Research Article
  • Citations3

Detecting Tweets Containing Cannabidiol-Related COVID-19 Misinformation Using Transformer Language Models and Warning Letters From Food and Drug Administration: Content Analysis and Identification.

  • Jan 23, 2023
  • JMIR Infodemiology
  • Jason Turner +3
  • Research Article
  • Citations27

Enhancing Transformer-based language models with commonsense representations for knowledge-driven machine comprehension

  • Mar 06, 2021
  • Knowledge-Based Systems
  • Ronghan Li +5
  • Video Transcripts

BERTAC: Enhancing Transformer-based Language Models with Adversarially Pretrained Convolutional Neural Networks

  • Aug 01, 2021
  • Underline Science Inc.
  • Kentaro Torisawa +3
  • Video Transcripts

An Ensemble of Transformer Based Models for Complex Named Entity Recognition Task

  • Jul 09, 2022
  • Underline Science Inc.
  • Fadi Hassan
  • Video Transcripts

Can Transformer Langauge Models Predict Psychometric Properties?

  • Jul 22, 2021
  • Underline Science Inc.
  • Animesh Nighojkar +3
  • PDF
  • Research Article
  • Citations18

Pre-trained transformer-based language models for Sundanese

  • Apr 13, 2022
  • Journal of Big Data
  • Wilson Wongso +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.