r/LocalLLaMA
· Communities
Local LLM open-source model options (5060TI 16GB)
Getting the above usage rate from running Qwen2.5-14B with the commands below ./llama-cli -m /home/XXXX/huggfacemodels/Qwen2.5-14B-Instruct-Q4_K_M.gguf -ngl 99 -c 32768 [ Prompt: 667.8 t/s | Generation: 44.0 t/s ] I think i can do better as there are still some headroom available on the gpu/cpu Any better way to get mo