Skip to content
X · @MoonshotAI · X / Twitter

We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. It delivers 1.72×–2.22× prefill spe…

We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels.It delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention.Explore on GitHub: http://github.com/MoonshotAI/FlashKDA