Skip to content
arXiv cs.AI · Papers

Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression

arXiv:2510.02345v4 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on dynamic expert clustering and structured compression to address these issues cohe