Skip to content
arXiv cs.CL · Papers

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

arXiv:2512.04013v3 Announce Type: replace Abstract: As augmented large language models (LLMs) with external tools become increasingly popular in web applications, improving augmented LLM inference serving efficiency and optimizing service-level objectives (SLOs) are critical for enhancing user experience. To achieve th