Skip to content
r/LocalLLaMA · Communities

Anyone running Deepseek v4 Flash with MoE offload?

I saw the DS4 repo and the last time I tried it I was just short of 5-10GB of VRAM to fit the model I wanted in VRAM with the KV cache. There are also these repos that caught my eye that I saw on the huihui-ai hugging face page - https://huggingface.co/huihui-ai/Huihui-DeepSeek-V4-Flash-abliterated-ds4-GGUF . The huihu