Skip to content
arXiv cs.CL · Papers

LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models

arXiv:2411.00918v5 Announce Type: replace Abstract: Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as GPT-OSS, DeepSeek-V3, Llama-4, and Gemini-2.5. However, systematic research on MoE remains severely constrained by the prohibi