ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving
In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE) models…