Skip to content
r/LocalLLaMA · Communities

Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization – AI’s narrative

# Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization I used Deepseek-v4-Flash-0731 cloud API settig up vllm-moet to run deepseek-v4-flash with MTP locally on single DGX Spark at 2-bit quant. Thought it might help others. Below is the summery from my AI Agent. So I did not write