arXiv cs.CL
· Papers
MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation
arXiv:2607.21978v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models remains underexplored. Existing PEFT methods for MoE either ignore router priors with uniform adapters, reducing efficiency