r/LocalLLaMA
· Communities
DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]
This is an an insane budget box that I've been using to test out a 98GB model using cpu generation on a 6 core CPU, 16gb vram.. for science. This week it went from 2t/s -> 7t/s on DeepSeek-V4-Flash-UD-Q2_K_XL, which has a 98GB vram requirement. Somewhere between b9986 and b10034 the llamacpp guys are cooking. [tequila