Skip to content
X · @teortaxesTex · X / Twitter

RT Red Hat AI: ~4x faster Kimi-K3 decoding. Our new DSpark speculator takes single-stream interactivity from ~110 to ~435 tok/s/user on math reasoning…

RT Red Hat AI~4x faster Kimi-K3 decoding.Our new DSpark speculator takes single-stream interactivity from ~110 to ~435 tok/s/user on math reasoning, delivers ~3.5x the output throughput at matched interactivity under load, and thanks to sliding window attention (2048-token window across all 5 draft layers), our model h