r/LocalLLaMA
· Communities
Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation
Ive been experimenting with task-aware GGUF quants for months, taking inspiration from TASA and TAQO but pushing the allocation lower to to the tensor level. The basic idea is to generate a custom imatrix from a category-specific corpus, measure where quantization causes damage, then redistribute a fixed bit budget tow