Skip to content
r/LocalLLaMA · Communities

How do you all decide which quantization format to use for deployment? (built a CLI to compare them, curious if others do this differently)

Every time I need to deploy a model I end up manually testing 3-4 quant formats to see what actually holds up on my hardware VRAM, speed, quality tradeoffs. Never found a clean way to compare them side by side instead of testing one at a time. Curious how others handle this do you just default to one format (AWQ/GPTQ s