Skip to content
r/LocalLLaMA · Communities

Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?

Hi. I don't have a GPU. So my biggest "local LLM" experience has been running ~26B models with single-digits tps values. However, the "serving economy" of DSv4 models look like a riddle to me. The Flash model has 284B parameters, but providers (e.g. OpenRouter) charge so little for it it's ridiculous. It's for example