r/LocalLLaMA
· Communities
40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)
daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster I would wager that compared to a naive kernel anyone can write it's more in the range of 10-20% faster e2e in reality, if at all, but hey, it's free and open! Apache 2.0 submitted by /u/Dany0 [link] [comme