Skip to content
r/LocalLLaMA · Communities

Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation

Ive been experimenting with task-aware GGUF quants for months, taking inspiration from TASA and TAQO but pushing the allocation lower to to the tensor level. The basic idea is to generate a custom imatrix from a category-specific corpus, measure where quantization causes damage, then redistribute a fixed bit budget tow