Skip to content
r/LocalLLaMA · Communities

Has anyone tested how quantization hits different capabilities separately? My results are surprising.

I've been running some systematic tests on a few models comparing FP16 vs various GGUF quant levels, and instead of looking at one aggregate benchmark score, I broke it down by capability: math (GSM8K), code (HumanEval), reasoning (ARC-Challenge), and knowledge recall (MMLU-Pro). The results are way more nuanced than "