arXiv cs.AI
· Papers
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
arXiv:2510.02345v4 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on dynamic expert clustering and structured compression to address these issues cohe