r/LocalLLaMA
ยท Communities
Deepseek v4 flash Q2 on a single 4090 ๐
It freaking worked lol๐ฅ Deepseek-v4-flash-0731 @UnslothAI 's IXQ2/Q3 checkpoint on one single RTX4090 with just 64 GB of RAM at usable token rate without dspark. All kernels running on Blaze (my custom developed ML compiler + inference engine) - no llama.cpp or vllm in the picture. The setup keeps heavily utilized expe