Skip to content
X · @teortaxesTex · X / Twitter

interestingly, I was wrong on who'll do this. I expected DeepSeek to go for 16 activated experts out of a large pool. Instead, they did 7/384, and Kim…

interestingly, I was wrong on who'll do this. I expected DeepSeek to go for 16 activated experts out of a large pool. Instead, they did 7/384, and Kimi 16/896. It was *not* obvious to everyone that you'd want to scale expert granularity.Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞): @dhtikna because 8/512 is what is used for 8