[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that:…