r/LocalLLaMA
· Communities
Dual Radeon AI PRO R9700 server much slower than RTX 5090 for LLM inference, Ollama bottleneck? vLLM / llama.cpp / other recommendations?
Hey All, First sorry for long post, and yes i dictated to AI and got it to fix grammer so its not painful for you all to read. Im trying to work out the optimal inference stack for a dedicated local AI server and would appreciate some advice from people running larger AMD/ROCm setups. Hardware Dedicated AI server: Ubun