Skip to content
r/LocalLLaMA · Communities

How does the Kv cache of MoEs scale?

I don't really understand how the Kv cache for MoE models scale. So like, if I take an example of a 35ba3b MoE, I know it uses the computation of a 3b+a bit more for routing, and the ram of 35b. But what about the kv cache? Does the ram needed for that scale as if its a 3b model or a 35b model? submitted by /u/Hot_Exam