any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?
I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization…
I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization…
I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload. My plan…
submitted by /u/InternationalGap3698 [link] [comments]
I made gemma4 12B write timestamp-anchored summaries of youtube video transcripts. I tested if the summaries have significant qualitative variance and if the SLM can pick…
Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC…
## Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q…
Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to…
Hi. So I just setup llamacpp for the first time. I'm using the model : "Huihui-Qwen3.6-35B-A3B-abliterated-ggml-model-Q4_K.gguf". When I test this in llamacpp server GUI I get…
Imagine you have 128gb of VRAM. what accompanying ram capacity you would choose (DDR4 8channel)? For example Deepseek v4 flash in q8 takes around 170GB +…
Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi…