Score: 0

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Published: December 3, 2025 | arXiv ID: 2512.04013v1

By: Ying Wang , Zhen Jin , Jiexiong Xu and more

Potential Business Impact:

Makes AI answer questions much faster.

Business Areas:

Semantic Search Internet Services

As augmented large language models (LLMs) with external tools become increasingly popular in web applications, improving augmented LLM inference serving efficiency and optimizing service-level objectives (SLOs) are critical for enhancing user experience. To achieve this, inference systems must maximize request handling within latency constraints, referred to as increasing effective throughput. However, existing systems face two major challenges: (i) reliance on first-come-first-served (FCFS) scheduling causes severe head-of-line blocking, leading to queuing delays exceeding the SLOs for many requests; and (ii) static batch token limit, which fails to adapt to fluctuating loads and hardware conditions. Both of these factors degrade effective throughput and service quality. This paper presents AugServe, an efficient inference framework designed to reduce queueing latency and enhance effective throughput for augmented LLM inference services. The core idea of AugServe is a two-stage adaptive request scheduling strategy. Specifically, AugServe combines the inference features of augmented LLM requests to optimize the order of scheduling decisions (stage I). These decisions are continuously refined with runtime information (stage II), adapting to both request characteristics and system capabilities. In addition, AugServe dynamically adjusts the token batching mechanism based on hardware status and real-time load, further enhancing throughput performance. Experimental results show that AugServe achieves 4.7-33.1x and 3.3-13.2x higher effective throughput than vLLM and InferCept, while reducing time-to-first-token (TTFT) by up to 96.3% and 95.0%, respectively.

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

Machine Learning (CS)

Makes AI answer questions much faster.

10 Apr 2025 2

89%

SpecServe: Efficient and SLO-Aware Large Language Model Serving with Adaptive Speculative Decoding

Computation and Language

Makes AI answer questions much faster.

7 Mar 2025 1

88%

Optimal Scheduling Algorithms for LLM Inference: Theory and Practice

Machine Learning (CS)

Makes AI answer questions much faster.

1 Aug 2025 0

View PDF Login to Bookmark

Page Count

15 pages

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Makes AI answer questions much faster.

Technical Abstract

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

SpecServe: Efficient and SLO-Aware Large Language Model Serving with Adaptive Speculative Decoding

Optimal Scheduling Algorithms for LLM Inference: Theory and Practice