Skip to content
arXiv cs.LG · Papers

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

arXiv:2607.08782v1 Announce Type: new Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing exper