Skip to content
r/LocalLLaMA · Communities

I tuned a custom Q8 for my AMD R9700s. My benchmark said 12.5% faster.

I have three Radeon AI PRO R9700s. I wanted to know if I could tune a quant for this specific hardware instead of just using the generic upstream formats. Came across the Github project for ROCmFPX by the wonderful Carlo Pasquale and started digging in. The fastest variant I've used is the MTP version. The format is Q8