X · @teortaxesTex
· X / Twitter
RT elie: the new sparse attention method introduced with this model is basically a combination of components from existing ones. let's go over each sp…
RT eliethe new sparse attention method introduced with this model is basically a combination of components from existing ones. let's go over each sparse attention method and what they keep from them:- deepseek sparse attention (DSA): they keep the top-k indexer, this is the basis of all the following sparse attention m