arXiv cs.LG
· Papers
Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
arXiv:2607.08782v1 Announce Type: new Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing exper