Skip to content
r/LocalLLaMA · Communities

Gemma 4 QAT could be improved further by Google aligning the QAT model to modern q4_k instead of q4_0

Hello, For the past few days I have been benchmarking Gemma 4 26b QAT UD Q4_K_XL extensively versus Bartowski's Q4_K_L. While QAT is certainly very effective and reducing memory consumption versus the highest q4 quant from him, I also have noticed some regressions in my own internal benchmarks I cannot share because I