r/LocalLLaMA
· Communities
Cursor releases their Mixture-of-Kittens megakernel for training MoE models – Claims to nearly double TFLOP/s
Link: https://cursor.com/blog/mixture-of-kittens GitHub: https://github.com/cursor/mixture-of-kittens Seems like a neat way to squeeze more performance out of MoE. I'm sure everyone has a favorite MoE model they'd like to try this with. It just dropped so I'm curious to hear people's opinions on it. submitted by /u/Cap