arXiv cs.CL
· Papers
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
arXiv:2512.04013v3 Announce Type: replace Abstract: As augmented large language models (LLMs) with external tools become increasingly popular in web applications, improving augmented LLM inference serving efficiency and optimizing service-level objectives (SLOs) are critical for enhancing user experience. To achieve th