X · @teortaxesTex
· X / Twitter
Not only does Gemma 4 31B have 2.4x more active parameters and more computationally intense attention But for each token of context, it uses 13.3x mor…
Not only does Gemma 4 31B have 2.4x more active parameters and more computationally intense attentionBut for each token of context, it uses 13.3x more bits of memory (at equal kv precision)As a result, at batch=32@256K seqlen, Gemma's cache is 320GB, more than DSV4 weights+kvdarren: how tf is deepseek flash cheaper tha