r/LocalLLaMA
· Communities
There’s a new PR for llamacpp claiming to boost prompt processing with rocm by around 15%, also fixes a bug which makes Q2_K 28x faster
This seems like a pretty solid improvement, and should make the more extreme quant setups viable on AMD cards. submitted by /u/Betadoggo_ [link] [comments]